πŸ€— HuggingFace Daily Papers
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
β–² 175 πŸŽ“ University of Technology Sydney
  • Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, en...
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
β–² 81 🏒 Salesforce AI Research
  • Stronger models can build inference-time harnesses that substantially improve weaker models' task performance without parameter updates by offloading ...
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
β–² 72 πŸŽ“ Zhejiang University
  • Mechanist is an autonomous agentic system that uses AI to discover and control the mechanisms underlying model intelligence, generating hypotheses, pe...
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
β–² 23 πŸŽ“ Beijing Jiaotong University
  • StateFlow introduces a persistent 3D world state to enable iterative, controllable previsualization for film and game design by constructing, evolving...
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent...
β–² 12 πŸŽ“ Chung-Ang University
  • Self-Geometry improves vision foundation model predictions by enforcing explicit multi-view geometric constraints via test-time adaptation with LoRA, ...
InSight-doc: Agentic Visual Perception for Long-Document Understanding
β–² 10 πŸŽ“ The Hong Kong University of Science and Technology
  • InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations.
Self-Evolving Embodied Agents via Skill-Harness Evolution
β–² 9 🏒 Microsoft Research
  • SHAPER is a train-free framework that improves embodied agents by evolving reusable skills and a context-code harness around a frozen foundation model...
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Vide...
β–² 6 🏒 Xiaomi Inc.
  • A closed-loop framework combining physics-based video synthesis, diffusion-based video dereflection, and a new benchmark achieves state-of-the-art vid...
Self-Refutation Loop 1Citation Validity 1Figure Editability 1Fabrication Detection 1Adversarial Review 1Composable Skills 1Integrity Checks 1Self-Critique 1Strong-To-Weak Scaffolding 1Inference-Time Harnesses 1Theory-Of-Mind Benchmarks 1Test-Time Capability Transfer 1
πŸ›οΈ Top Research Institutions
Improved cross-validated distances for multivariate pattern analysis
πŸ›οΈ UniversitΓ© de MontrΓ©al AI & Machine Learning
  • Language models often struggle with alignment to demographic groups, leading to biased or unrepresentative outputs.
  • The paper explores group alignment-induced sycophancy, proposing a two-sided evaluation framework for steerable pluralis...
The Fallacy of Independent Ceilings: Characterizing Coupled Load-Branch Stall Interaction
πŸ›οΈ University of Rhode Island Systems & Infrastructure
  • The need for efficient and scalable document extraction from large collections of documents is a significant challenge i...
  • Scout proposes a scalable method for document extraction using data similarity techniques to enhance processing efficien...
Welfare Approximation in Multilateral Trade
πŸ›οΈ Tel Aviv University Theory & Algorithms
  • Mechanism design in multilateral trade is complex due to the requirement for unanimous agreement among multiple agents.
  • Introduction of a new mechanism that facilitates efficient trade by considering the interactions among all agents involv...
VeriFin: A Neurosymbolic Framework for Verifying LLM-Generated Financial Claims
πŸ›οΈ Stevens Institute of Technology Other CS
  • Verifying numerical claims generated by large language models is challenging due to potential inaccuracies in reporting ...
  • A neurosymbolic framework that combines symbolic reasoning with neural network capabilities to verify LLM-generated fina...
Optimal Experimental Design and Estimation when Potential Outcomes are Bounded
πŸ›οΈ OpenAI Economics
  • Designing randomized experiments to estimate treatment effects is often complicated by bounded potential outcomes.
  • A method for optimal experimental design and analysis that accounts for bounded potential outcomes in randomized experim...
AI & Machine Learning
418 papers 7 cats
Systems & Infrastructure
29 papers 6 cats
Software & Programming
15 papers 4 cats
Theory & Algorithms
29 papers 5 cats
Applications
75 papers 7 cats
Other CS
32 papers 9 cats
Quantitative Finance
9 papers 9 cats
Economics
7 papers 3 cats