#research
Every summary, chronological. Filter by category, tag, or source from the rail.
DRSR: Reducing Long-Horizon Agent Compute via Deletion Risk
DRSR (Deletion Risk for Set-level Representation) optimizes long-horizon AI agents by identifying and pruning redundant or low-utility information from the agent's memory set, significantly reducing compute overhead without sacrificing task performance.
TimeEvo: Improving Time Series Agents via Failure-Driven Evolution
TimeEvo enhances time series forecasting agents by implementing a self-evolution loop that analyzes past failures to iteratively refine reasoning strategies and model performance.
Policy-as-Skill: Deterministic Governance for LLM Decision Support
The 'Policy-as-Skill' framework integrates deterministic governance into LLM workflows by treating organizational policies as executable skills, ensuring decisions are evidence-based, auditable, and constrained by hard rules.
Provably Complete Generalized Planning with LLMs
This research introduces a framework for achieving provably complete generalized planning using LLMs, moving beyond heuristic-based generation to ensure reliable, verifiable task execution across diverse problem instances.
Optimizing Small Language Models with Minimum Risk Training
Minimum Risk Training (MRT) significantly improves the performance of small language models in specialized tasks like power outage report generation by optimizing for task-specific metrics rather than standard cross-entropy loss.
Predicting Objective Conflict in Pluralistic AI Alignment
This research introduces a framework for identifying which AI objectives are inherently conflicting, allowing developers to implement 'dials' for steerable, pluralistic alignment rather than forcing a single, static optimization path.
Introducing MentalHealthBench: Evaluating AI in Mental Health
OpenAI has released MentalHealthBench, an open-source evaluation framework developed with over 80 global mental health experts to measure how AI models handle realistic, non-emergency and acute mental health conversations.
Evaluating Input Representations for Multimodal Document QA
This research evaluates whether multimodal document QA models perform better using raw pixel data, extracted text, or a hybrid approach, finding that representation choice significantly impacts accuracy and efficiency.
The Linear Representation Hypothesis in Neural Networks
The Linear Representation Hypothesis posits that neural networks encode complex, high-dimensional concepts as linear directions within their internal activation spaces, allowing for simple geometric manipulation of model outputs.
MAWILE: A Multi-Axis Workbench for Evaluating LLM Evaluators
MAWILE provides a structured framework to audit and inspect LLM-based evaluators, addressing the critical need to validate the reliability of automated evaluation systems.
LLM Judge Consensus Often Overstates Accuracy Due to Error Dependence
Using multiple LLM judges to reach a consensus does not guarantee higher accuracy because these models share systematic error dependencies, leading to inflated confidence in incorrect outputs.
EvidenT: Grounding Enterprise AI in Evidence and Traceability
EvidenT is a framework designed to improve the reliability of enterprise AI assistants by enforcing strict evidence grounding and providing verifiable traceability for every generated response.
Optimizing Medical LLMs: Didactic Knowledge vs. Clinical Cases
The paper investigates how different data types—structured didactic knowledge versus unstructured clinical case reports—impact the reasoning and diagnostic capabilities of medical LLMs.
Establishing Global Standards for Frontier AI and RSI
To safely navigate the acceleration of AI research and recursive self-improvement (RSI), the industry must move toward shared international technical standards for safety, evaluation, and incident reporting.
Efficient Production Benchmarking for LLM Agents
Static benchmarks are insufficient for production LLM agents; continuous evaluation using real-world historical data is required to track performance as models and user inputs evolve.
Implicit Rule Induction via Test-Time Task Embeddings
This paper introduces a method for solving ARC-like reasoning tasks by generating test-time task embeddings that implicitly capture underlying transformation rules, enabling models to generalize to novel patterns without explicit rule programming.
The AI-GRACE Framework for Operationalizing Agentic AI
AI-GRACE is a structured framework designed to bridge the gap between high-level organizational goals and the technical architecture required to deploy reliable, compliant agentic AI systems.
SpecOpt: Agentic Molecule Optimization via Contact-Diff Reasoning
SpecOpt introduces a novel agentic framework for molecular optimization that uses 'Contact-Diff' reasoning to improve binding specificity, moving beyond simple affinity metrics to address complex protein-ligand interactions.
Clinician-Grounded QA for AI-Assisted Psychiatric Intake
This research proposes a framework for quality assurance in AI-assisted psychiatric intake by grounding AI outputs in clinical standards, ensuring safety and accuracy in sensitive mental health assessments.
CogGym: Benchmarking Human vs. Machine Cognition at Scale
CogGym provides a standardized framework for comparing AI model performance against human cognitive benchmarks, addressing the need for rigorous, large-scale evaluation of machine intelligence.
TinyCeNN-LM: Efficient Model Compression via Cellular-Recurrent Layers
TinyCeNN-LM introduces a method to replace standard attention mechanisms in pretrained LLMs with Cellular Neural Network (CeNN)-inspired recurrent layers, significantly reducing computational overhead while maintaining performance through quality-gated conversion.
Detecting LLM Hallucinations via Topological Context Analysis
This research proposes a method to detect LLM hallucinations by identifying topological signatures of 'impaired context sharing' within the model's internal activations, offering a structural approach to reliability.
Decoupling Internal Representations from Causal Importance in LLMs
Fine-tuning often causes significant shifts in internal model representations that do not necessarily correlate with causal importance, suggesting that model behavior changes are localized in specific, sparse components rather than global weight updates.
A Unified Evaluation Framework for Trustworthy AI Systems
The provided source is a placeholder for an academic paper on evaluating LLMs, agents, and multimodal systems. It highlights the industry's shift toward standardized, trustworthy benchmarks for complex AI architectures.
Self-Improvement via Fast Tree-Search
The paper presents a methodology for enhancing AI model performance by integrating fast tree-search algorithms, enabling models to iteratively improve their reasoning and output quality through structured exploration.
LLM-as-an-Improver: Iterative Candidate Refinement
Instead of using LLMs only for verification, the 'LLM-as-an-Improver' framework uses feedback from verifiers to iteratively refine and improve candidate outputs, significantly increasing success rates in complex reasoning tasks.
Risks of Agent-Mediated Hiring: Access and Recurrence Bias
Multi-agent résumé screening systems can inadvertently amplify hiring biases, creating 'recurrence' where specific candidate profiles are consistently excluded due to agent-to-agent feedback loops.
Characterizing Web Search by Conversational LLM Agents
This research analyzes how conversational LLM agents navigate web search, identifying distinct search strategies and the relationship between query formulation, search results, and final response quality.
Mapping the Design of LLM Benchmarks
Current LLM benchmarks often lack transparency in design, leading to misaligned evaluations. This research provides a taxonomy to categorize benchmark construction, helping developers better understand what models are actually being tested for.
Detecting LLM Harm via Latent States
Rather than relying on output filtering, this research proposes monitoring internal latent states of LLMs to detect harmful intent before it manifests in generated text.
Showing 30 of 482