Building Reliable Long-Horizon Agents: A Survey
Preprint, 2026
A survey of the definitions, metrics, benchmarks, and system-design principles needed to build agents that remain reliable over extended tasks.
Preprint, 2026
A survey of the definitions, metrics, benchmarks, and system-design principles needed to build agents that remain reliable over extended tasks.
Under review at EACL 2027, 2026
A framework for automated code-efficiency optimization that mines reusable agent skills from slow and optimized program pairs and applies them through execution-free diagnosis and rewriting.