#computer-vision
Every summary, chronological. Filter by category, tag, or source from the rail.
Cooperative Multi-Agent Driving via V2V-VLA Models
CMU-Drive introduces a reasoning-focused benchmark for multi-agent autonomous driving, while V2V-VLA enables vehicles to share visual and linguistic insights to improve collective decision-making.
Adversarially Robust Abductive Fusion for Perception Models
This paper introduces a framework for combining pre-trained transformer perception models using abductive reasoning to improve robustness against adversarial attacks.
Building a Durable Memory Layer for Video Intelligence
Video AI systems fail because they treat video as a bag of frames rather than a spatial-temporal volume. To build true video memory, you must ingest once, store primitives like entities and relationships in a context graph, and ground every reasoning step in timestamps.
AI EngineerCOMPASS: Improving Compositional Control in Multimodal Models
COMPASS introduces a unified framework that uses a shared 'expert token' to bridge composition perception and generation, enabling precise layout control in multimodal models.
OmniPath: Automating Wheelchair Accessibility Audits with AI
OmniPath improves accessibility mapping by fusing OpenStreetMap data with high-density LiDAR to identify physical barriers like slope and surface discontinuities that standard maps ignore.
SpatialClaw: Using Code as an Action Interface for Spatial Reasoning
SpatialClaw is a training-free agent framework that improves spatial reasoning in VLMs by treating Python code—rather than structured tool calls—as the primary interface for perception and geometric tasks.
Qwen-RobotSuite: Three Foundation Models for Embodied AI
The Qwen team has released a suite of three specialized foundation models—RobotManip, RobotWorld, and RobotNav—designed to address data fragmentation in robotics through unified action representations, language-conditioned world modeling, and scalable navigation interfaces.
Edge-Based Computer Vision for Industrial Food Waste Reduction
Mill uses custom-tuned Gemma models on Nvidia Jetson hardware to process high-frame-rate video at the edge, turning food waste data into actionable procurement insights for commercial kitchens.
Google Cloud TechByteDance's Lance: A Unified 3B Model for Vision and Video
Lance is an open-source, 3B parameter unified model that natively integrates image and video understanding, generation, and editing within a single jointly trained framework.
Showing 9 of 9