AI Research

Latest research papers from arXiv covering machine learning, computer vision, natural language processing, and more.

arXivPDF

Robot Learning with Visual Predicted Force

Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimat...

Haonan Chen, Feiyang Wu, Yuxiang Ma
Oct 3, 2026
arXivPDF

PyINE: A Framework for Scalable Elicitation and Oversight via Code Execution

Reasoning models can remain capable of solving a task while still defaulting to cheaper but misleading shortcuts. This creates a central oversight problem: when a model gives an answer with plausible but incomplete reasoning, can an overseer determine whether that output should be trusted? To study ...

Pierre-Luc St-Charles, Alessandro Palmas, Damiano Fornasiere
Oct 3, 2026
arXivPDF

Verb-ICL: Rethinking In-Context Learning for Structured Prediction

Structured prediction tasks pose unique challenges for in-context learning (ICL): their compositional outputs require modeling fine-grained, token-level patterns that sentence-level approaches fail to capture, and their task-specific annotation conventions are human-defined artifacts that cannot be ...

Fan Bai, Hengshuo Miao, Sanjit S Batra
Oct 3, 2026
arXivPDF

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis a...

Ramil Khafizov, Ilya Statsenko, Ruslan Rakhimov
Oct 3, 2026
arXivPDF

WNet: Discrete Wavelets Transform for Efficient Token Mixing

In a Transformer, token mixing is the step that lets each token draw information from other tokens, and it dominates the cost of encoding long sequences. Self-attention does this mixing very well: every token weighs every other token by content, which gives strong contextual modeling. That all-pairs...

Rana Aref Salama, Abdou Youssef, Mona Diab
Oct 3, 2026
arXivPDF

Learning to Clarify Underspecified Intents Under Limited Interaction

AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information problem: the assistant should acquire i...

Pranav M R, Manuel Cherep, Pattie Maes
Oct 3, 2026
arXivPDF

Towards Automatically Pruning Logging Code with Coding Agents: How Far Are We?

Logging code supports debugging, monitoring, and software maintenance, but excessive logging can add noise, impose runtime overhead, and obscure diagnostic information. While prior research has extensively studied logging code generation and modification, logging removal remains comparatively undere...

He Yang Yuan, Haonan Zhang, Xin Wang
Oct 3, 2026
arXivPDF

GPU-Accelerated Bregman Douglas-Rachford Splitting for Discrete Optimal Transport

We present GPU-accelerated Bregman Douglas--Rachford splitting algorithm (BDRS) for discrete optimal transport problem in three input formats: an explicit cost matrix, a point cloud with a ground cost between them, and a separable cost on a regular grid. For each input format, we propose hardware-aw...

Yifan Xu, Shiqian Ma
Oct 3, 2026

Data from arXiv.org • Updated hourly