#research
Every summary, chronological. Filter by category, tag, or source from the rail.
Building Whistleblowing Infrastructure for AI Agents
New reporting tools allow AI agents to flag misbehaving peers, but experts warn that fostering collaboration through positive models is more effective than building an automated surveillance state.
The Andrej Karpathy Blog: A Decade of AI Engineering
Andrej Karpathy's blog serves as a foundational archive of practical AI engineering, emphasizing 'from-scratch' implementations, deep learning fundamentals, and the importance of hands-on experimentation.
MOSAIC: Query-Aware Exploration for GraphRAG
MOSAIC improves GraphRAG performance by dynamically adapting exploration policies based on the specific query, moving beyond static traversal methods to retrieve more relevant graph-based context.
KuaiRP: Technical Report on Role-Playing Model Optimization
The KuaiRP technical report details specialized training methodologies for enhancing LLM performance in role-playing scenarios, focusing on character consistency and narrative depth.
The Agent Incident Registry: A Framework for Preventing AI Failures
The Agent Incident Registry (AIR) proposes a standardized, community-driven database to catalog and analyze AI agent failures, enabling developers to learn from past errors and prevent recurring systemic vulnerabilities.
Quantifying the Memorization-to-Generalization Transition in Grokking
The paper provides a quantitative framework for understanding 'grokking'—the phenomenon where neural networks suddenly shift from memorizing training data to generalizing—by identifying specific scaling laws and phase transitions in model learning.
Training Nemotron for Olympiad-Level Mathematics
The paper outlines a systematic recipe for training LLMs to achieve gold-medal performance in Olympiad-level mathematics, emphasizing high-quality synthetic data generation and iterative reinforcement learning.
Deterministic Math Solvers for Clinical LLMs
To address the unreliability of LLMs in clinical settings, this paper proposes a deterministic math solver architecture that separates reasoning from calculation, ensuring accuracy in high-stakes medical computations.
Beyond Task Completion: Measuring AI Agent Resilience
Current AI agent benchmarks focus too heavily on final success, ignoring 'resilience'—the ability to maintain performance and considerate behavior under mounting environmental pressure.
Task-Agnostic Environment Preprocessing for AI Agents
The paper introduces a method for AI agents to learn from environments without predefined task syllabi, focusing on task-agnostic preprocessing to improve generalization and performance.
Calibrating AI Agent Confidence via Internal Representations
AI agents often struggle to self-evaluate success. This research proposes a method to calibrate confidence by analyzing internal model representations rather than relying on external feedback or output text.
CityPlanner: A Sandbox Agent for Executable Urban Planning
CityPlanner is an AI agent framework designed to simulate urban planning by executing plans within a sandbox environment, allowing for iterative refinement and evaluation of complex city development strategies.
Operational Architecture for Cognitive Digital Twins
The paper proposes a shift from simple state synchronization in digital twins to an architecture enabling cognitive self-evolution, allowing systems to learn and adapt autonomously.
Automated Black-Box Red Teaming for Agentic AI Systems
A systematic framework for identifying risks in agentic AI by using a taxonomy-driven approach to automate black-box red teaming, moving beyond manual testing to discover vulnerabilities in complex, multi-step agent workflows.
PRAGMA: Enhancing Long-Term AI Memory Alignment
PRAGMA introduces a framework for evaluating how well AI models maintain personalized, consistent guidance across lifelong conversations by measuring memory alignment.
Evaluating AI Scientist Workflows with OpenDiscoveryTrace
OpenDiscoveryTrace provides a standardized dataset and framework for evaluating the multi-step reasoning and discovery processes of AI agents acting as scientists.
Subagents vs. Agent Skills for Long-Horizon Tasks
The article evaluates architectural patterns for complex AI workflows, comparing the modularity of subagents against the efficiency of reusable agent skills in executing long-horizon tasks.
Evaluating Explainable AI (XAI) Quality with LLMs
The XAI-Arena framework tests whether Large Language Models can reliably evaluate the quality of explainable AI outputs, aiming to automate the subjective process of human-centric XAI assessment.
EnvCraft: Automating Environment Synthesis for Agentic RL
EnvCraft introduces a framework for synthesizing executable environments to train 'claw-like' agents, addressing the bottleneck of manual environment design in reinforcement learning.
Optimizing Business Processes with Control-Flow Uncertainty
This paper introduces a mathematical framework for scheduling business processes where the execution path is uncertain, using stochastic optimization to balance resource allocation and process completion time.
Decision-Targeted Evaluation for Human-Agent Teams
Traditional AI evaluation metrics fail to capture the nuances of human-agent collaboration. This paper proposes a decision-targeted framework that prioritizes the quality of final outcomes over individual task completion.
Why AI Agents Over-Trust Unreliable Tools
AI agents frequently fail to verify tool outputs, leading to 'tool-reliance bias' where models blindly accept incorrect data from external APIs or functions.
CriticGen: Improving LLM Evaluation via Generation-Aware Feedback
CriticGen shifts AI evaluation from static scoring to an iterative, generation-aware process, providing actionable feedback that directly improves model performance by aligning critiques with the specific generation context.
Protecting Reasoning Circuits During LLM Compression
Standard model compression often degrades reasoning performance by pruning critical model weights; 'Reasoning-Aware Compression' identifies and protects these specific circuits to maintain intelligence while reducing energy consumption.
Budget-Aware Online Adaptation for Web Agents
This paper introduces a framework for optimizing web agent performance by selectively teaching models only when necessary, balancing the high cost of fine-tuning against the gains in task completion accuracy.
When Does Memory Help? A Cost-Aware Evaluation of LLM Agents
Long-term memory in LLM agents often introduces diminishing returns; performance gains frequently fail to justify the increased latency and token costs unless the task requires high-fidelity historical context.
SciLitBench: Evaluating LLMs for Systematic Literature Reviews
SciLitBench provides a standardized benchmark and design framework for evaluating how LLMs perform in systematic literature reviews, identifying critical gaps in reasoning and evidence extraction for scientific research.
Automating Quantum Chip Calibration with AI Agents
Integrating AI agents with laboratory software allows researchers to automate routine, multi-step quantum chip calibration, shifting human effort from manual monitoring to high-level experimental design.
Rethinking Aleatoric Uncertainty in LLMs via Interpretations
The paper argues that LLM uncertainty in ambiguous tasks stems from multiple valid interpretations rather than simple randomness, proposing a shift from answer-based to interpretation-based uncertainty estimation.
BioSync: Transformer-Based Cross-Modal Fusion for Digital Biomarkers
BioSync introduces a transformer-based architecture designed to fuse disparate physiological data streams into a unified digital biomarker, improving predictive accuracy in multimodal health monitoring.
Showing 30 of 438