Observatorio IA
Lo último en inteligencia artificial, con fuente y fecha
Radar diario de investigación publicada en arXiv, comparador de proveedores y calendario del Reglamento europeo de IA.
Contenido elaborado con ayuda de IA y revisado por nuestro equipo editorial. Los resúmenes son extractos breves; consulta siempre la fuente original.
Pregunta al Observatorio
Radar de investigación
- cs.RO· Jusuk Lee, Sungha Kim, Yeonsoo Park et al.
Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video.
Leer en arXiv - cs.CV· Ritesh Thawkar, Shubham Patle, Shravan Venkatraman et al.
Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models. Such supervision is costly to obtain and can reward plausible failures: a realistic output may leave the requested change undone or alter content that should…
Leer en arXiv - cs.RO· Junyan Li, Ruizhi Li, Yu Liu et al.
DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias p…
Leer en arXiv - cs.RO· Lizhi Yang, Yiling Hou, Yao Tang et al.
CSF: Contextual Safety Filtering for Motion Generators
Text-conditioned motion generators produce trackable whole-body motion, but they have no notion of scene-dependent safety: the same action may target an object or a person. Existing safeguards either inspect the prompt, require labeled motion data, or enforce geometric constraints; therefore, they do not directly ac…
Leer en arXiv - cs.AI· Drew T. Nguyen, William Fithian
On the estimation and validity of AI time horizons---a statistical look at the METR plot
METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI di…
Leer en arXiv - cs.RO· Octi Zhang, Mateo Guaman Castro, Patrick Yin et al.
A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations.
Leer en arXiv - cs.CR· Abbas Raftari
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope. The paths were different.
Leer en arXiv - cs.CV· Jiahua Dong, Anurag Bagchi, Yash Jangir et al.
What 30,000 Hours of Ego-centric Video Does Not Teach
World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors.
Leer en arXiv - cs.CV· You-Zhe Xie, Ting-Wei Chou, Yu-Hsuan Li et al.
OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint.
Leer en arXiv - cs.CV· Ankan Deria, Komal Kumar, Hisham Cholakkal et al.
WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete.
Leer en arXiv - cs.CV· Zhongyu Yang, Jiale Tao, Ruitao Chen et al.
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide cov…
Leer en arXiv - cs.HC· Nhan, Tran, Neal Wadhwa et al.
Hybrid Cinematography: Previsualizing and Managing Hallucination Risk in Generative Video Reshooting
On a film set, the camera move is committed during a take. Generative video reshooting lets filmmakers change it afterward, but may require hallucinating unrecorded content, a gap sometimes discovered only after leaving the set.
Leer en arXiv - cs.AI· Peter Kulits, Yiqing Xu, R. Kenny Jones et al.
BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built.
Leer en arXiv - cs.CV· Boyao Han, Chen Shi, Jingjing Qian et al.
VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable fr…
Leer en arXiv - cs.CV· Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combinat…
Leer en arXiv - cs.LG· Anna Zimmel, Fleur Hendriks, Markus Holzleitner et al.
Bi-FORK: Generative Modeling of High-Dimensional Bifurcating Systems
Bifurcations are ubiquitous in physical systems, from structural buckling to fluid and climate dynamics, yet they remain largely unexplored in deep learning. At a symmetry-breaking bifurcation, a single input admits multiple equally valid solutions, violating the one-to-one assumption underlying most learned physica…
Leer en arXiv - cs.LG· Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba et al.
Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel pr…
Leer en arXiv - cs.LG· Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al.
Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses…
Leer en arXiv - cs.CV· Suhwan Cho, Yonwoo Choi, Soongjin Kim et al.
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point clo…
Leer en arXiv - stat.ML· Song Liu
Density Ratio Estimation with Stein Displacement Fields
Density ratios quantify distribution shift from a probability-mass point of view, whereas displacement fields describe, from a dynamical point of view, how one distribution is transported onto another. Although both offer complementary insights, they are usually estimated separately, and converting one into the othe…
Leer en arXiv
Comparador de proveedores
Precio orientativo: verifica siempre en la web del proveedor. Enlaces a sus páginas oficiales de precios.
| Proveedor | Modelos | Modalidades | Precios oficiales |
|---|---|---|---|
| OpenAI | Familia GPT | Texto, imagen, audio | Ver precios |
| Anthropic | Familia Claude | Texto, imagen | Ver precios |
| Familia Gemini | Texto, imagen, audio, vídeo | Ver precios | |
| Mistral AI | Mistral / Codestral | Texto | Ver precios |
| Meta | Llama (pesos abiertos) | Texto, imagen | Ver precios |
Calendario del Reglamento de IA de la UE
1 ago 2024
Entrada en vigor del Reglamento (UE) 2024/1689 de IA
2 feb 2025
Prohibiciones de prácticas de IA y obligación de alfabetización en IA
2 ago 2025
Obligaciones para modelos de IA de uso general y gobernanza
2 ago 2026
Aplicación general, incluidos sistemas de alto riesgo del Anexo III (calendario sujeto a revisión normativa)
2 ago 2027
Sistemas de alto riesgo integrados en productos del Anexo I
