πŸ€— HuggingFace Daily Papers
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
β–² 25 πŸŽ“ HKU MMLab
  • UniClawBench introduces a capability-driven benchmark for evaluating proactive agents in real-world environments using live Docker container evaluatio
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Ge
β–² 25 πŸŽ“ Shanghai Jiao Tong University
  • A benchmark for scientific lineage reasoning and idea generation is introduced, organizing scientific works as genetic-like Idea Genome objects and ev
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolatio
β–² 22 πŸŽ“ National Yang Ming Chiao Tung University
  • LongE2V enables high-quality video recovery from sparse event streams by leveraging pre-trained video diffusion priors and addressing temporal stabili
DrugGen 2: A disease-aware language model for enhancing drug discovery
β–² 14 πŸŽ“ Isfahan University of Medical Sciences
  • DrugGen-2 generates small molecules conditioned on disease ontology and target protein sequences through fine-tuning GPT-2 with supervised learning an
Enhancing In-context Panoramic Generation via Geometric-aware Pretraining
β–² 14 🏒 Insta360 Research
  • Canvas360 is a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with fine-tuning, featuring a large-sc
OpenCoF: Learning to Reason Through Video Generation
β–² 9 🏒 ByteDance
  • OpenCoF framework introduces a reasoning video dataset and model that improve temporal reasoning through diverse supervision and explicit reasoning to
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Genera
β–² 6 🏒 Meituan
  • OPSD-V enhances few-step autoregressive video diffusion models by using real long-video data for temporal context during training, providing dense tra
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
β–² 5 🏒 Meta AI
  • In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act.
  • As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed
  • We call this failure mode "behavioral state decay".
  • We study memory as an active intervention mechanism rather than passive retrieval.
Large Language Models 1Multimodal Large Language Models 1Proactive Agents 1Real-World Environments 1Capability-Driven Benchmark 1Skill Usage 1Exploration 1Long-Context Reasoning 1Multimodal Understanding 1Cross-Platform Coordination 1Docker Containers 1Closed-Loop Evaluation 1
πŸ›οΈ Top Research Institutions
Closed-form fractional radial links for elliptical Mahalanobis discriminant analysis
πŸ›οΈ Cherkasy State Business College AI & Machine Learning
  • The need for interpretable machine learning models is increasingly critical as their applications expand across various
  • A method for steering neural network training through interpretable constraints based on partial dependence.
Validity of LLMs as data annotators: AMALIA on authority
πŸ›οΈ Universidade LusΓ³fona AI & Machine Learning
  • Large language models (LLMs) often struggle with reasoning and logical consistency, leading to unreliable outputs.
  • A framework for quantifying uncertainty, coherence, and robustness of LLMs using a graph-based approach.
Steering Neural Network Training through Interpretable Constraints Based on Partial Depend
πŸ›οΈ University of LiΓ¨ge AI & Machine Learning
  • The efficiency of model evaluation is hindered by fixed-size benchmarks that do not adapt to diverse evaluation objectiv
  • A method for efficient model evaluation that determines the optimal amount of data needed for testing.
GradInf: Gradient Estimation as Probabilistic Inference
πŸ›οΈ Carnegie Mellon University Software & Programming
  • Gradient estimation in probabilistic programs is notoriously difficult and has diverse applications in scientific comput
  • A novel method for gradient estimation framed as probabilistic inference.
Locality of Curve-Decoding and Improved Proximity Gaps
πŸ›οΈ Massachusetts Institute of Technology Theory & Algorithms
  • Understanding the locality of curve-decoding in error-correcting codes is crucial for improving decoding efficiency.
  • A study on proximity gaps in the context of Interactive Oracle Proofs (IOPs) and Succinct Non-interactive Arguments of Z
Sculptable Mesh Structures for Room-Scale Form-Finding
πŸ›οΈ Carnegie Mellon University Applications
  • Designing physical structures effectively within the constraints of a computer interface is challenging.
  • A sculptable mesh structure approach that facilitates room-scale form-finding through interactive design.
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
πŸ›οΈ Renmin University of China Other CS
  • Singing voice synthesis (SVS) has limited generalization across diverse musical genres, affecting its practical applicab
  • A benchmarking framework, MMGenre, designed to evaluate SVS performance across multiple musical genres.
Estimating the Stochastic Discount Factor from Option Prices and Predicting the Equity Pre
πŸ›οΈ The University of Tokyo Quantitative Finance
  • Estimating the stochastic discount factor (SDF) from option prices is complex and often inaccurate.
  • A new framework that scales the SDF by time-varying volatility using market data from S&P 500 options.
Estimating Causal Effects from Data Generated by Stochastic Algorithms
πŸ›οΈ Stanford University Economics
  • Estimating causal effects from data generated by stochastic algorithms presents unique challenges.
  • The paper proposes a method for estimating causal effects in contexts where content is selected based on user characteri
AI & Machine Learning
155 papers 7 cats
Systems & Infrastructure
29 papers 6 cats
Software & Programming
19 papers 4 cats
Theory & Algorithms
38 papers 5 cats
Applications
83 papers 7 cats
Other CS
19 papers 8 cats
Quantitative Finance
15 papers 9 cats
Economics
8 papers 3 cats