#agents
Every summary, chronological. Filter by category, tag, or source from the rail.
Governed Persistent Memory for Long-Horizon AI Agents
This research introduces a 'Governed Persistent Memory' framework that uses source-bound state semantics and fail-closed release mechanisms to improve reliability and safety in long-horizon AI agents.
MindMemOS: A Self-Evolving Memory Layer for AI Agents
MindMemOS introduces a portable, self-evolving memory operating layer that decouples agent intelligence from long-term storage, enabling persistent, adaptive memory across diverse AI architectures.
AstraZeneca's Agentic R&D Research Assistant
AstraZeneca has developed an agentic AI system designed to automate complex R&D workflows, demonstrating how large-scale pharmaceutical research can leverage autonomous agents to accelerate discovery.
Moving AI Agents from Game-Based RL to Real-World Reliability
Training AI agents for computer use requires moving beyond simple outcome-based reinforcement learning toward 'flight school' simulations that account for real-world messiness, partial observability, and adversarial UI.
AI EngineerPractical Loop Engineering for AI Agents
Loop engineering uses autonomous feedback cycles to automate repetitive tasks. By combining 'goal' primitives for bounded tasks and 'loop' primitives for scheduling, developers can build reliable agentic workflows while maintaining human oversight for critical judgment.
Industrial AI Scaling, Local Models, and Cybersecurity Risks
The panel discusses the shift toward industrial-scale AI infrastructure, the rise of high-performance local models like Meta's Muse Glimmer, and the emerging cybersecurity implications of autonomous agent capabilities in upcoming models like OpenAI's Astra.
Optimizing Agentic Workflows with GPT-5.6
GPT-5.6 shifts the economics of agentic AI by enabling high-performance results with smaller models, reduced reasoning effort, and new API primitives like programmatic tool calling and multi-agent orchestration.
Building Lifelong AI Research Partners via Agent Memory
To transform AI from a stateless tool into a lifelong research partner, systems must implement persistent, context-aware memory architectures that allow agents to retain domain-specific knowledge and evolve alongside materials scientists.
Automating Process Engineering Diagrams with LLMs
This research explores a multi-agent framework for generating and validating Process Flow Diagrams (PFDs) and Piping and Instrumentation Diagrams (P&IDs) using LLMs to reduce manual engineering errors.
Optimizing AI Harnesses to Slash Enterprise Token Costs
Writer’s new Palmyra X6 model and upgraded agentic harness aim to reduce enterprise AI costs by up to 50% by focusing on infrastructure efficiency rather than just model selection.
Google 'All Things Agentic' Hackathon Overview
Google is hosting a global hackathon with $180,000 in prizes, challenging developers to build autonomous, production-ready AI agents using Gemini 3.5 and Google Cloud.
Moving from AI Assistance to Agentic Execution
Enterprise AI is shifting from Q&A to autonomous execution. 'Frontier firms'—the top 10% of users—are outpacing others by 8.3x in output volume by integrating agents with company-specific tools, data, and repeatable workflows.
Sparse Coding for Latent Communication in VLM Agents
This paper introduces a post-hoc sparse coding method to interpret and analyze the latent communication signals exchanged between vision-language model (VLM) agents, providing a framework for understanding multi-agent internal states.
Evaluation-Conditioned Training for Stronger Oversight
Evaluation-Conditioned Training (ECT) improves model performance by training agents to adapt their behavior based on the strength of the oversight regime they operate under, ensuring better generalization.
SBCO: Self-Supervised Verifier-Grounded Harness Optimization
SBCO is a framework for optimizing planning agents by using self-supervised, verifier-grounded harness optimization to improve decision-making accuracy without requiring extensive human-labeled data.
MESA: Task-Adaptive Evidence Selection for Agent Memory
MESA improves long-horizon agent performance by using a task-adaptive, multi-structure memory selection framework that retrieves relevant evidence more effectively than standard retrieval methods.
Automating Behavioral Research for AI Agents
This paper introduces a framework for scaling behavioral scientific research on AI agents, moving beyond manual evaluation to automated, reproducible experimental pipelines.
Continuously Improving AI Agents via Trace Data Mining
To improve autonomous agents, treat them like machine learning models: collect execution traces, mine them for feedback, and use that data to iteratively refine prompts, fine-tune models, and update agent state.
AI EngineerThe Evolution and Future of AI Memory Systems
AI memory has evolved from simple thread-based context to asynchronous 'running profiles.' Despite this, current systems suffer from staleness, lack of cross-platform integration, and an inability to reason over external data sources like email or calendars.
Raising the Floor: Practical AI Agent Evaluation
Stop chasing benchmark scores and start treating agent evaluations like production software tests. Focus on identifying when issues start and their impact on user volume to build reliable, trust-based AI products.
Continual Learning via Distillation: A 2x2 Taxonomy
Enterprises can implement continual learning by mapping distillation tasks across a 2x2 grid of offline/online traces and hints, allowing for immediate performance improvements without requiring 'golden' datasets.
Building an Automated LLM-Powered Knowledge Base
Transform disorganized raw notes into a structured, interconnected wiki using voice dictation, LLM-based enrichment, and automated cloud-based pipelines.
Democratizing Frontier AI: Automating Discovery and Scaling
The era of massive, monolithic pre-training is hitting a ceiling. By automating model training and data optimization, we can shift the focus from compute-heavy scaling to domain-specific innovation, allowing more builders to participate at the frontier.
Scaling Expertise: Moving Beyond Raw Intelligence in AI Agents
Current AI agents excel at symbolic tasks like coding but struggle with real-world digital work because they lack 'expertise'—the ability to learn and adapt to idiosyncratic micro-worlds through continuous learning.
Scaling Compute on Context: Moving Beyond Public Data
Current AI models excel on public data but fail to acquire deep, personalized knowledge. The solution lies in 'scaling compute on context'—using recursive self-improvement to deepen a model's understanding of private data without hitting a synthetic data wall.
Building Memory Harnesses for Long-Horizon AI Agents
To prevent context rot in long-horizon AI tasks, implement a structured 'write-manage-read' memory loop. A ranked recall policy consistently outperforms basic RAG or no-memory baselines, improving accuracy while reducing token costs.
Scaling Continual Learning with On-Policy Self-Distillation
On-Policy Self-Distillation (OPSD) enables models to learn continuously from real-world data by using privileged hints to guide training, overcoming the infrastructure and reward-density limitations of traditional RLHF and GRPO.
Moving Beyond Checklists: Operationalizing AI and SBOM Security
Security experts argue that frameworks like the OWASP Top 10 and SBOM guidance are not compliance checklists but foundations for cyber resilience, requiring active tabletop exercises and operational integration to be effective.
Agent-MD: Automating Scientific Simulations with LLM Orchestration
Agent-MD introduces a framework for stateful Grand Canonical Monte Carlo (GCMC) and Molecular Dynamics (MD) campaigns, using selective LLM intervention and event-driven escalation to manage complex simulation workflows autonomously.
Preventing AI Research Drift with Structured Scientific Loops
To prevent AI research agents from drifting, researchers must enforce 'scientific taste' and falsifiable constraints within the automated loop, specifically applied here to quadruped navigation.
Showing 30 of 1278