#agents
Every summary, chronological. Filter by category, tag, or source from the rail.
TimeEvo: Improving Time Series Agents via Failure-Driven Evolution
TimeEvo enhances time series forecasting agents by implementing a self-evolution loop that analyzes past failures to iteratively refine reasoning strategies and model performance.
Policy-as-Skill: Deterministic Governance for LLM Decision Support
The 'Policy-as-Skill' framework integrates deterministic governance into LLM workflows by treating organizational policies as executable skills, ensuring decisions are evidence-based, auditable, and constrained by hard rules.
Provably Complete Generalized Planning with LLMs
This research introduces a framework for achieving provably complete generalized planning using LLMs, moving beyond heuristic-based generation to ensure reliable, verifiable task execution across diverse problem instances.
Improving LLM Agent Training with Subtask Decomposition
RLDS improves agent training by replacing scalar trajectory rewards with subtask-specific advantage estimation, allowing models to learn more effectively from complex, multi-step tasks.
Decoupling Proposal and Judgment in AI-Driven Investment Research
To prevent false discoveries in AI-driven factor mining, researchers must separate the agent's proposal role from a frozen, anytime-valid statistical referee that judges performance based solely on future market outcomes.
JAZ: A Minimalist Agent Framework Using Code as a Harness
JAZ replaces complex, specialized agent harnesses with a single 'invoke' primitive, allowing LLMs to manage memory and self-improvement through recursive code execution.
Identifying Silent Failures in AI Agent-Tool Interactions
AI agents often suffer from 'silent failures' where tool invocations appear successful but return incomplete or incorrect data, silently propagating errors downstream into final outputs.
ElevenLabs Strategy: Scaling Voice AI and Enterprise Adoption
ElevenLabs is scaling to $600M ARR by positioning its voice models as a critical enterprise layer, prioritizing market share over immediate margins, and focusing on emotional intelligence to pass the Turing test.
Scaling Autonomous Drone Fleets as Infrastructure
Skydio is shifting drone operations from manual piloting to autonomous, agentic infrastructure by splitting intelligence between edge-based flight safety and cloud-based VLM orchestration.
Building Agent-Native Communication Platforms
Ando is a team messaging platform that treats AI agents as first-class participants rather than external integrations, aiming to eliminate the 'meat proxy' bottleneck where humans manually relay information between agents and teams.
Using AI Agents and APIs for Real-Time Data Processing
LLMs are poor at raw data crunching but excellent at reasoning. By offloading heavy computation to specialized APIs and using AI agents to orchestrate tool-calling, you can ground models in real-time, high-fidelity data.
Scaling AI Engineering: How Airbnb Integrates GPT-6 Astra
Airbnb has expanded its partnership with OpenAI to integrate GPT-6 Astra across its product and engineering teams, moving beyond code generation into system design, debugging, and marketplace operations.
Self-Organizing Agent Teams Learn Collaborative Reasoning
Self-Organizing Agent Teams (SAT) move beyond fixed AI workflows by learning reusable strategies for role division and information flow, enabling teams to solve complex problems that individual agents cannot handle alone.
EvidenT: Grounding Enterprise AI in Evidence and Traceability
EvidenT is a framework designed to improve the reliability of enterprise AI assistants by enforcing strict evidence grounding and providing verifiable traceability for every generated response.
Meta's Muse Agent Strategy: Scaling via Ecosystem Integration
Meta is aggressively expanding its Muse AI agent by integrating it into hardware (smart glasses), desktop OS (macOS), and third-party commerce platforms, aiming to monetize through transaction fees rather than subscription models.
Building Embodied Foundation Models with Perceptive Objectives
Perceptron AI is moving beyond traditional VLMs by unifying perception, reasoning, and control into a single 'embodied foundation model' that uses data-sparse mixture-of-experts to handle context bloat and learns task-relevant percepts automatically.
AI EngineerImplementing Real-Time Tool Calling for Voice AI Agents
To build responsive voice agents, use a synchronous tool-calling loop where the model decides and your code executes. Keep tools instant to avoid conversation gaps and use a 'before_tool_callback' checkpoint to enforce policies and handle slow actions.
Building Reliable AI Agents: The Data-First Approach
Moving from RAG to agentic workflows requires treating document processing as a multi-step pipeline where data quality, structured representation, and agentic harnesses are critical to preventing compounding errors.
Scaling Multi-Agent Video Analysis at Meta
Meta manages 100M+ videos using a specialized multi-agent pipeline that detects modality misalignment and unoriginal content through domain-specific VLMs, continuous DPO, and aggressive compute optimizations.
Building the Document Context Layer for AI Agents
Modern RAG is shifting from simple retrieval to agentic workflows where document parsing, semantic storage, and specialized extraction pipelines act as the critical context layer for autonomous agents.
Efficient Production Benchmarking for LLM Agents
Static benchmarks are insufficient for production LLM agents; continuous evaluation using real-world historical data is required to track performance as models and user inputs evolve.
The AI-GRACE Framework for Operationalizing Agentic AI
AI-GRACE is a structured framework designed to bridge the gap between high-level organizational goals and the technical architecture required to deploy reliable, compliant agentic AI systems.
SpecOpt: Agentic Molecule Optimization via Contact-Diff Reasoning
SpecOpt introduces a novel agentic framework for molecular optimization that uses 'Contact-Diff' reasoning to improve binding specificity, moving beyond simple affinity metrics to address complex protein-ligand interactions.
The Dark Arts of Skill Engineering
Moving beyond basic prompting, skill engineering treats AI as a harness extension. By using adversarial sub-agents, deterministic linters, and external scripts to force divergence, you can escape the 'median gravity' of model outputs and build truly robust AI tools.
AI EngineerMoving Beyond Token Consumption to Outcome-Based AI
Measuring AI success by token consumption leads to either wasteful 'tokenmaxxing' or counterproductive 'token minimization.' Organizations should instead adopt 'valuemaxxing'—a strategy that prioritizes measurable operational outcomes like deployment speed and rework reduction over raw usage volume.
Optimizing Inference for Agentic Workflows
Agentic inference requires shifting focus from individual request latency to end-to-end task completion, utilizing prefix caching and agent-aware scheduling to reduce costs and improve performance.
AI EngineerAdvances in Data Center Inference Engineering
Inference engineering is shifting from post-training optimization to a cycle where dedicated training processes—specifically in quantization, KV compaction, and speculative decoding—are essential for production performance.
Operating Distributed Inference Systems at Scale
Inference at scale is no longer a model problem; it is an orchestration problem. Reliability and efficiency now depend on a unified control plane that manages GPU state, KV cache, and distributed request routing.
A Unified Evaluation Framework for Trustworthy AI Systems
The provided source is a placeholder for an academic paper on evaluating LLMs, agents, and multimodal systems. It highlights the industry's shift toward standardized, trustworthy benchmarks for complex AI architectures.
Architecting Long-Horizon AI Agents via Cascaded Intelligence
The paper proposes a hierarchical architecture for long-horizon AI agents that decouples high-level strategic planning from low-level execution using 'levels' and 'ticks' to manage complex, multi-step tasks.
Showing 30 of 1531