#ai-agents
Every summary, chronological. Filter by category, tag, or source from the rail.
Querying and Acting on Cloud Data with Data Agent Kit
The Data Agent Kit provides a unified framework of MCP servers, agent skills, and IDE integrations that allow AI agents to securely query, analyze, and modify data across BigQuery, Cloud SQL, and Cloud Storage.
Google Cloud TechThe Economics of Web Context: Renting vs. Owning for AI Agents
For high-frequency AI knowledge work, renting context via APIs becomes prohibitively expensive. Building an owned data pipeline often reaches a cost-efficiency tipping point at surprisingly low volumes (around 15,000 queries).
AI EngineerFixing Computer Use Benchmarks: Beyond Replay Exploits
Current computer use benchmarks are often gamed by 'replay agents' that blindly repeat successful trajectories. Robust evaluation requires stochastic, verified environments and honest statistical uncertainty to avoid costly deployment errors.
Why Computer-Use Models Will Agentify the Web
The web was built for human eyes, not APIs. Instead of waiting for a universal API layer, AI agents will 'agentify' the web by interacting directly with pixels and DOMs, treating browsers as game engines to perform tasks.
Synchronizing Beliefs via Second-Order Theory-of-Mind
This paper proposes a framework for human-autonomy teams where agents model human beliefs about the agent's own state to reduce misalignment and improve collaborative performance.
The CASE Framework for Enterprise Agentic AI Governance
The CASE Framework provides a multi-disciplinary architecture to govern enterprise AI agents by integrating technical, legal, and operational controls into a unified oversight structure.
Gary Tan on Founder Psychology, AI Agency, and First Principles
Gary Tan discusses the evolution of Silicon Valley, the importance of founder earnestness over trend-chasing, and how AI agents are fundamentally changing the speed and scale of building.
a16z (Andreessen Horowitz)MetaSpace: Metamorphic Testing for Embodied AI Spatial Cognition
MetaSpace introduces a metamorphic testing framework to evaluate spatial reasoning in embodied agents by applying geometric transformations to environments and verifying if agent behavior remains consistent.
KNOWPLAN: Knowledge-Driven AI Agents for Degree Planning
KNOWPLAN is an AI agent framework that integrates structured knowledge graphs with LLMs to solve complex academic degree pathway planning, ensuring adherence to institutional constraints and student goals.
Running AI Agents in Production Without the On-Call Tax
Engineering teams spend 70% of their time on operational overhead rather than coding. By deploying autonomous background agents that leverage production context, teams can automate incident triage, deployment monitoring, and routine operational tasks, effectively offloading the 'on-call tax'.
AI EngineerRefactoring Legacy Codebases in the Age of AI Agents
While AI models are rapidly improving, they cannot yet reliably 'one-shot' complex refactors. Building a clean, maintainable monorepo remains a high-ROI investment that accelerates development velocity and improves developer experience.
AI EngineerProject2Task: Graph-Guided Planning for Autonomous Research
Project2Task improves autonomous research agents by using graph-based planning to decompose high-level project goals into actionable, structured task sequences, overcoming the limitations of linear prompt-based planning.
Beyond Agents: Building AI-Native Software
Agents are the 'web pages' of our era—a primitive, not the destination. The next frontier is AI-native software that leverages asynchronous context, dynamic interfaces, and multi-agent orchestration.
AI EngineerCloudflare Launches Kitesurf: A Headless Browser for AI Agents
Cloudflare has introduced Kitesurf, a cloud-hosted, headless browser built on Workers, designed specifically for AI agents to navigate the web efficiently without the overhead of traditional consumer browsers.
SafeCommit: Certifying Safety for Memory-Grounded AI Agents
SafeCommit is a framework that introduces a certification mechanism to determine when memory-grounded AI agents can safely execute actions based on their internal state, reducing the risk of hallucinated or harmful operations.
Naïve Raises $28.5M to Automate Autonomous Business Operations
Naïve provides an API-first infrastructure that allows AI agents to provision and manage business operations—from incorporation to cloud resources—while building specialized runtime layers to reduce the high costs of agent inference.
Gadgets: Personal AI-Driven App Development on Cloudflare
Kenton Varda introduces 'Gadgets,' a platform where AI agents can safely modify and extend individual app instances, bypassing traditional plugin architecture bottlenecks by leveraging isolated, container-free infrastructure.
AI EngineerLocalizing AI Agent Failures: Model vs. Harness
To debug AI agents effectively, you must distinguish between failures caused by the underlying LLM (Model) and those caused by the agent's orchestration, tools, or environment (Harness).
Agentic Engineering: From Writing Code to Orchestrating Systems
Agentic engineering shifts the developer's role from writing deterministic code to designing, constraining, and supervising autonomous AI systems that operate on probabilistic judgment.
IBM TechnologyDesigning AI Agents to Minimize Hallucination
AI agents hallucinate because they are trained to prioritize fluent, confident pattern completion over factual accuracy. You can mitigate this by grounding agents in real-time data, enforcing tool-based verification, strictly defining operational scope, and implementing human-in-the-loop oversight.
IBM TechnologyClinLens: Long-Horizon Coding Agents for Clinical Data Science
ClinLens is an AI agent framework designed to handle the complexities of longitudinal, multimodal clinical data by automating long-horizon coding tasks in data science workflows.
Automating Healthcare Administration with AI Agents
Lassie is replacing manual administrative labor in healthcare practices with AI agents that handle billing, insurance, and scheduling, allowing providers to focus on patient care rather than paperwork.
a16z (Andreessen Horowitz)ProcAgent: Edge-Based Procedural Guidance with Human-in-the-Loop
ProcAgent is an agentic framework designed to provide real-time, procedural task guidance on edge devices by integrating human-in-the-loop feedback to improve accuracy and reliability in complex workflows.
Integrating AI Agents into Event-Sourced Systems
Improve fraud detection by layering agentic AI onto existing event-sourced architectures, using a semantic layer to provide agents with the necessary context to resolve ambiguous transactions.
Building Autonomous Software Factories with Forward Deployed Engineering
Forward deployed engineering is shifting from manual consulting to building 'software factories'—autonomous systems where AI agents handle the full lifecycle from signal to deployment, provided the codebase is 'agent-ready' with robust validation loops.
AI EngineerScaling Forward Deployed Engineering with Scoping and AI Agents
Forward Deployed Engineering (FDE) requires balancing rigorous manual scoping to avoid 'feature bloat' with the automation of repetitive pipeline tasks using AI agents to maintain competitive velocity.
Automating Performance Engineering with AI Agents at Netflix
Netflix uses AI agents to bridge the gap between profiling data and production code fixes, creating a self-improving catalog of performance anti-patterns that allows for automated, canary-validated optimizations.
Applying Control Theory to AI Coding Agents
Instead of using AI agents to generate massive, unreviewable pull requests, use control theory to build iterative loops that make small, verifiable, and incremental code changes.
AI EngineerAutomating Incident Response with Self-Improving Agents
Observability is shifting from passive dashboards to active telemetry for AI agents. By feeding production traces directly into code-aware sandboxes, teams can automate root cause analysis and generate pull requests for fixes.
AI EngineerWhy Cognition Acquired Poke: The Shift Toward AI Personality
Cognition, the maker of Devin, acquired AI assistant startup Poke to integrate its conversational, personality-driven interaction model into their coding agent, signaling that user experience and 'colleague-like' rapport are becoming key competitive advantages.
Showing 30 of 140