#llm
Every summary, chronological. Filter by category, tag, or source from the rail.
Policy-as-Skill: Deterministic Governance for LLM Decision Support
The 'Policy-as-Skill' framework integrates deterministic governance into LLM workflows by treating organizational policies as executable skills, ensuring decisions are evidence-based, auditable, and constrained by hard rules.
Improving AI Agent Robustness Against Incentive-Misaligned Environments
Computer-use agents often fail to act in a user's best interest when environments are designed to steer outcomes. The CAVEAT benchmark reveals that performance drops from 78.6% to 17.3% under steering, but targeted interventions can recover 55% of that performance.
Provably Complete Generalized Planning with LLMs
This research introduces a framework for achieving provably complete generalized planning using LLMs, moving beyond heuristic-based generation to ensure reliable, verifiable task execution across diverse problem instances.
Optimizing Small Language Models with Minimum Risk Training
Minimum Risk Training (MRT) significantly improves the performance of small language models in specialized tasks like power outage report generation by optimizing for task-specific metrics rather than standard cross-entropy loss.
Improving LLM Agent Training with Subtask Decomposition
RLDS improves agent training by replacing scalar trajectory rewards with subtask-specific advantage estimation, allowing models to learn more effectively from complex, multi-step tasks.
JAZ: A Minimalist Agent Framework Using Code as a Harness
JAZ replaces complex, specialized agent harnesses with a single 'invoke' primitive, allowing LLMs to manage memory and self-improvement through recursive code execution.
Identifying Silent Failures in AI Agent-Tool Interactions
AI agents often suffer from 'silent failures' where tool invocations appear successful but return incomplete or incorrect data, silently propagating errors downstream into final outputs.
TwinCheck: Verifying Stateful AI Agents via Negative-Twin Simulation
TwinCheck improves agent reliability by creating 'negative twins'—simulated environments that test if an agent's proposed action leads to unintended state changes before execution.
Building Production-Ready Apps with Gemini 3.5 Transcribe
Gemini 3.5 Transcribe offers two distinct APIs for speech-to-text: synchronous batch processing for pre-recorded files and the Live API for real-time streaming, both supporting advanced features like diarization, word-level timestamps, and custom vocabulary.
Google Cloud TechBuilding AI-Powered Transcription Pipelines with Gemini 3.5
Gemini 3.5 Transcribe enables developers to build high-accuracy, domain-specific transcription pipelines for both live and batch audio without requiring model training.
Using AI Agents and APIs for Real-Time Data Processing
LLMs are poor at raw data crunching but excellent at reasoning. By offloading heavy computation to specialized APIs and using AI agents to orchestrate tool-calling, you can ground models in real-time, high-fidelity data.
Scaling AI Agents: How Ringg Achieves 65% Call Resolution
Ringg uses a multi-model OpenAI orchestration layer to automate customer service, achieving 65% resolution rates and 90% cost reductions by routing tasks to specialized models.
How Invideo Uses GPT-6 Astra for Agentic Video Editing
Invideo leverages GPT-6 Astra to automate complex video editing tasks, achieving a 3x improvement in color-grading success rates and enabling the rapid creation of custom, editable effects.
The Linear Representation Hypothesis in Neural Networks
The Linear Representation Hypothesis posits that neural networks encode complex, high-dimensional concepts as linear directions within their internal activation spaces, allowing for simple geometric manipulation of model outputs.
MAWILE: A Multi-Axis Workbench for Evaluating LLM Evaluators
MAWILE provides a structured framework to audit and inspect LLM-based evaluators, addressing the critical need to validate the reliability of automated evaluation systems.
LLM Judge Consensus Often Overstates Accuracy Due to Error Dependence
Using multiple LLM judges to reach a consensus does not guarantee higher accuracy because these models share systematic error dependencies, leading to inflated confidence in incorrect outputs.
Optimizing Medical LLMs: Didactic Knowledge vs. Clinical Cases
The paper investigates how different data types—structured didactic knowledge versus unstructured clinical case reports—impact the reasoning and diagnostic capabilities of medical LLMs.
Building Embodied Foundation Models with Perceptive Objectives
Perceptron AI is moving beyond traditional VLMs by unifying perception, reasoning, and control into a single 'embodied foundation model' that uses data-sparse mixture-of-experts to handle context bloat and learns task-relevant percepts automatically.
AI EngineerImplementing Real-Time Tool Calling for Voice AI Agents
To build responsive voice agents, use a synchronous tool-calling loop where the model decides and your code executes. Keep tools instant to avoid conversation gaps and use a 'before_tool_callback' checkpoint to enforce policies and handle slow actions.
Scaling Regenerative Agriculture with AI-Driven Pasture Management
Labor-intensive rotational grazing is the primary barrier to sustainable livestock farming. By using AI agents to analyze environmental data and automate decision-making for virtual fencing, we can scale pasture-based systems to compete with industrial feedlots.
Stop Deploying VLMs: Use Vibe Training for Task-Specific Models
Avoid deploying Vision Language Models (VLMs) at runtime due to latency and licensing issues. Instead, use a 'vibe training' pipeline: leverage VLMs to auto-label datasets, use ensemble judges to filter quality, and train small, Apache 2.0-licensed models like RF-DETR for production-grade performance.
Building the Document Context Layer for AI Agents
Modern RAG is shifting from simple retrieval to agentic workflows where document parsing, semantic storage, and specialized extraction pipelines act as the critical context layer for autonomous agents.
Optimizing GPT-6 Prompt Caching for Persistent Agents
OpenAI has updated GPT-6 with improved prompt caching, offering up to 90% discounts on cached tokens and new diagnostic tools to monitor hit rates, diagnose misses, and optimize context reuse for long-running agents.
OpenAI Launches GPT-6 Sol and Luna with 50% Price Reductions
OpenAI has expanded the GPT-6 family with Sol and Luna, two cost-efficient models that bring Astra-level intelligence to professional workflows, coding, and computer use at half the price of their predecessors.
Building Institutional Memory with V7's Context Graph
V7 Go uses a structured 'Context Graph' to turn scattered enterprise data into persistent, queryable memory for AI agents, enabling complex, multi-step workflows with high accuracy and auditability.
Efficient Production Benchmarking for LLM Agents
Static benchmarks are insufficient for production LLM agents; continuous evaluation using real-world historical data is required to track performance as models and user inputs evolve.
TinyCeNN-LM: Efficient Model Compression via Cellular-Recurrent Layers
TinyCeNN-LM introduces a method to replace standard attention mechanisms in pretrained LLMs with Cellular Neural Network (CeNN)-inspired recurrent layers, significantly reducing computational overhead while maintaining performance through quality-gated conversion.
Detecting LLM Hallucinations via Topological Context Analysis
This research proposes a method to detect LLM hallucinations by identifying topological signatures of 'impaired context sharing' within the model's internal activations, offering a structural approach to reliability.
Decoupling Internal Representations from Causal Importance in LLMs
Fine-tuning often causes significant shifts in internal model representations that do not necessarily correlate with causal importance, suggesting that model behavior changes are localized in specific, sparse components rather than global weight updates.
Optimizing Transformer Inference with FlashNorm
FlashNorm accelerates transformer inference by folding RMS norm gains into projection weights and parallelizing normalization and matrix multiplication via custom CUDA kernels.
AI EngineerShowing 30 of 1401