AI & LLMs
The deepest channel on Edge. Foundation models, agent architectures, retrieval, evals, and the moving line between research and production.
DRSR: Reducing Long-Horizon Agent Compute via Deletion Risk
DRSR (Deletion Risk for Set-level Representation) optimizes long-horizon AI agents by identifying and pruning redundant or low-utility information from the agent's memory set, significantly reducing compute overhead without sacrificing task performance.
TimeEvo: Improving Time Series Agents via Failure-Driven Evolution
TimeEvo enhances time series forecasting agents by implementing a self-evolution loop that analyzes past failures to iteratively refine reasoning strategies and model performance.
Policy-as-Skill: Deterministic Governance for LLM Decision Support
The 'Policy-as-Skill' framework integrates deterministic governance into LLM workflows by treating organizational policies as executable skills, ensuring decisions are evidence-based, auditable, and constrained by hard rules.
Improving AI Agent Robustness Against Incentive-Misaligned Environments
Computer-use agents often fail to act in a user's best interest when environments are designed to steer outcomes. The CAVEAT benchmark reveals that performance drops from 78.6% to 17.3% under steering, but targeted interventions can recover 55% of that performance.
Provably Complete Generalized Planning with LLMs
This research introduces a framework for achieving provably complete generalized planning using LLMs, moving beyond heuristic-based generation to ensure reliable, verifiable task execution across diverse problem instances.
Optimizing Small Language Models with Minimum Risk Training
Minimum Risk Training (MRT) significantly improves the performance of small language models in specialized tasks like power outage report generation by optimizing for task-specific metrics rather than standard cross-entropy loss.
Improving LLM Agent Training with Subtask Decomposition
RLDS improves agent training by replacing scalar trajectory rewards with subtask-specific advantage estimation, allowing models to learn more effectively from complex, multi-step tasks.
Decoupling Proposal and Judgment in AI-Driven Investment Research
To prevent false discoveries in AI-driven factor mining, researchers must separate the agent's proposal role from a frozen, anytime-valid statistical referee that judges performance based solely on future market outcomes.
Predicting Objective Conflict in Pluralistic AI Alignment
This research introduces a framework for identifying which AI objectives are inherently conflicting, allowing developers to implement 'dials' for steerable, pluralistic alignment rather than forcing a single, static optimization path.
JAZ: A Minimalist Agent Framework Using Code as a Harness
JAZ replaces complex, specialized agent harnesses with a single 'invoke' primitive, allowing LLMs to manage memory and self-improvement through recursive code execution.
Identifying Silent Failures in AI Agent-Tool Interactions
AI agents often suffer from 'silent failures' where tool invocations appear successful but return incomplete or incorrect data, silently propagating errors downstream into final outputs.
TwinCheck: Verifying Stateful AI Agents via Negative-Twin Simulation
TwinCheck improves agent reliability by creating 'negative twins'—simulated environments that test if an agent's proposed action leads to unintended state changes before execution.
Building Production-Ready Apps with Gemini 3.5 Transcribe
Gemini 3.5 Transcribe offers two distinct APIs for speech-to-text: synchronous batch processing for pre-recorded files and the Live API for real-time streaming, both supporting advanced features like diarization, word-level timestamps, and custom vocabulary.
Google Cloud TechCustomizing Flux: From Generative Media to Robotics
Black Forest Labs demonstrates how to extend foundational video models like Flux beyond creative media into action prediction and robotics through prompt upsampling, modular moderation, and weight-based fine-tuning.
Building Reliable Generalist Robots via Active Learning
Dyna Robotics achieves 99.4% reliability in complex tasks like napkin folding by using reward models to detect failures, enabling targeted active learning and error recovery rather than relying on massive, uncurated datasets.
Building Embodied AI: Why World Models Need Causality
Christopher Manning argues that current generative video models are insufficient for robotics because they lack underlying semantics. Moonlake AI is building action-conditioned world models that allow for physical interaction and planning, aiming to replace 10,000 hours of teleoperation with simulation.
Using AI Agents and APIs for Real-Time Data Processing
LLMs are poor at raw data crunching but excellent at reasoning. By offloading heavy computation to specialized APIs and using AI agents to orchestrate tool-calling, you can ground models in real-time, high-fidelity data.
Scaling AI Engineering: How Airbnb Integrates GPT-6 Astra
Airbnb has expanded its partnership with OpenAI to integrate GPT-6 Astra across its product and engineering teams, moving beyond code generation into system design, debugging, and marketplace operations.
Introducing MentalHealthBench: Evaluating AI in Mental Health
OpenAI has released MentalHealthBench, an open-source evaluation framework developed with over 80 global mental health experts to measure how AI models handle realistic, non-emergency and acute mental health conversations.
Scaling AI Literacy Through Community-Led Training
After two years and 4 million engagements, OpenAI Academy is shifting from direct facilitation to a 'train-the-trainer' model, empowering local organizations to lead their own practical AI workshops.
Scaling Practical AI Literacy for Gig Economy Workers
OpenAI and Grab are launching 'GO Forward with AI,' a two-year training program designed to teach 30,000 gig workers and merchants in Southeast Asia how to apply AI tools to business planning, sales analysis, and operations.
How Invideo Uses GPT-6 Astra for Agentic Video Editing
Invideo leverages GPT-6 Astra to automate complex video editing tasks, achieving a 3x improvement in color-grading success rates and enabling the rapid creation of custom, editable effects.
Evaluating Input Representations for Multimodal Document QA
This research evaluates whether multimodal document QA models perform better using raw pixel data, extracted text, or a hybrid approach, finding that representation choice significantly impacts accuracy and efficiency.
The Linear Representation Hypothesis in Neural Networks
The Linear Representation Hypothesis posits that neural networks encode complex, high-dimensional concepts as linear directions within their internal activation spaces, allowing for simple geometric manipulation of model outputs.
MAWILE: A Multi-Axis Workbench for Evaluating LLM Evaluators
MAWILE provides a structured framework to audit and inspect LLM-based evaluators, addressing the critical need to validate the reliability of automated evaluation systems.
Self-Organizing Agent Teams Learn Collaborative Reasoning
Self-Organizing Agent Teams (SAT) move beyond fixed AI workflows by learning reusable strategies for role division and information flow, enabling teams to solve complex problems that individual agents cannot handle alone.
LLM Judge Consensus Often Overstates Accuracy Due to Error Dependence
Using multiple LLM judges to reach a consensus does not guarantee higher accuracy because these models share systematic error dependencies, leading to inflated confidence in incorrect outputs.
EvidenT: Grounding Enterprise AI in Evidence and Traceability
EvidenT is a framework designed to improve the reliability of enterprise AI assistants by enforcing strict evidence grounding and providing verifiable traceability for every generated response.
Optimizing Medical LLMs: Didactic Knowledge vs. Clinical Cases
The paper investigates how different data types—structured didactic knowledge versus unstructured clinical case reports—impact the reasoning and diagnostic capabilities of medical LLMs.
Meta's Muse Agent Strategy: Scaling via Ecosystem Integration
Meta is aggressively expanding its Muse AI agent by integrating it into hardware (smart glasses), desktop OS (macOS), and third-party commerce platforms, aiming to monetize through transaction fees rather than subscription models.
Showing 30 of 1605