LeVJEPA is introduced, the first video encoder trained under LeJEPA's collapse-free objective, which indicates that video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.
Highlighted research
Patch Policy: Efficient Embodied Control via Dense Visual Representations
Patch Policy provides evidence that high-performance embodied intelligence may emerge from modular systems that combine powerful pretrained representations with small, specialized control policies.
Read the highlight →HighlightedInterpreting Physics in Video World Models
One of the first interpretability analyses of how video encoders (V-JEPA 2) represent physical variables.
Read the highlight →This work introduces LpWorldModel, a JEPA model regularized with Rectified Distribution Matching Regularization to match encoder features to a Rectified Generalized Gaussian distribution, yielding non-negative sparse codes, and finds that the learned sparse representations are mode-factored.
RayDINO, a large-scale self-supervised visual encoder for chest X-rays trained on 840,000 images and evaluated on 82,000 images sourced from 12 publicly available datasets, is introduced, suggesting that self-supervision enables patient-centric analysis useful in clinical workflows, allowing for a more holistic interpretation of X-rays.
This first release of Prior Labs in relational learning opens-source three pieces of software that expect to accelerate research in the field towards meaningful real-world impact and releases an initial version of RPI, an open-source, model-agnostic interface that enables early adopters to easily define problems on new databases and apply any model implemented in RelArena, including TabPFN-Rel, to these problems.
A multi-agent collaborative framework, StarVerus, to automate the verification of industrial Rust code and introduces a planner-repairer-actor-rewriter multi-agent paradigm to further enhance the proof repair capabilities.
This study provides empirical clarity through a systematic exploration of multimodal pretraining and derives efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget.
It is shown that self-supervised learning (SSL) can convert raw DXA images into representations of systemic health, and prognostic information in DXA images that conventional readouts discard is revealed, learnable with relatively little data and modest compute.
Horsepower-JEPA organizes each graph into an ordered bank of coarse-to-fine partition resolutions and performs context-target latent prediction separately at each resolution using an online encoder, an exponential-moving-average target encoder, and a latent predictor.
A framework for analyzing openness across the AI stack is proposed, reviewing prior approaches and highlighting the diverse motivations for pursuing openness, showing how these map onto different forms of openness at both the model and system levels.
This work proposes to learn a world model of piano sound using JEPA by framing music as an action-conditioned system, and shows that the learned model captures the relationships between musical actions and their resulting sound.
This work proposes CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction, and demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.
This work proposes Anchored Self-Play (ASP), which anchors self-play with a small reference set by adding a code-embedding similarity reward for generation and mixing reference bugs into fixer training, and achieves the best fix rates across bug sources.
Off‐the‐shelf LLMs are not yet suitable for autonomous requirement quality evaluation; yet could be more suitable to assist human experts or serve specialized roles in agentic AI frameworks.
This study evaluates the performance of lightweight LLMs (under 8B parameters) in generating biomedical abstracts and lay summaries in a zero-shot setting and introduces a novel analysis of the sectional origin and desirability of information.
C-JEPA is a simple and flexible object-centric world model that extends masked joint embedding prediction from image patches to object-centric representations and induces latent interventions with counterfactual-like effects and prevents shortcut solutions, making interaction reasoning essential.
AdaJEPA is an adaptive latent world model that performs test-time adaptation within the closed loop of model predictive control (MPC), and substantially improves planning success with as few as one gradient step per MPC replanning step.
This work introduces MJEPA, a joint-embedding predictive architecture for audio-visual learning that uses a single, unified encoder for both modalities, and shows that cross-modal prediction is critical: without it, a shared encoder degrades below unimodal baselines; with it, each modality's representation benefits from the other.
The JEPA-style model for real-time quadrotor control is introduced, which combines a latent dynamics model with a novel physics-inspired prober that maps frozen latents to interpretable state, enabling physically grounded long-horizon prediction.
This work provides the first provable analysis of successful internalization: for the task of learning parities, it is shown that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, learning to directly compute the parity.
S-JEPA is introduced, a JEPA-style encoder-predictor pair trained to match the soft posteriors of a Gaussian Mixture Model at masked positions via KL divergence, establishing a new Pareto frontier without offline re-clustering or teacher distillation.
Hyperball is a simple optimizer wrapper that improves learning rate transfer across widths and depths compared to decoupled weight decay, motivated by prior theory showing that training with weight decay leads to an equilibrium weight norm that only depends on the training hyperparameters.
Temporal Difference in Vision (TDV) is introduced, a new paradigm for self-supervised learning from video that avoids existing inductive biases, relying instead on a causal assumption that the past causes the future.
Evaluated across several robotics benchmarks, WorldDP consistently outperforms existing baselines, validating that coupling the world model's physically grounded planning with diffusion policy's efficient execution yields superior multi-stage performance.
The trained Llama-4 system substantially improves objective intervention quality over strong proprietary baselines and open-weight baselines across all six datasets, and Oracle-plan experiments further show that, when plan quality is controlled, the trained duplex model produces high-quality guidance and large gains on Out-of-Plan recovery.
It is proved that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions, and the Gaussian is the unique latent distribution for which this guarantee holds.
These findings demonstrate that PGT effectively address the bottleneck of fine-grained perception, revealing that many spatial reasoning deficits stem from inadequate supervision signals rather than inherent architectural or resolution limitations.
Cambrian-P is revisited as a lightweight supervisory signal and Cambrian-P, a video MLLM augmented with per-frame learnable camera tokens and a pose regression head is introduced, which achieves substantial gains on spatial reasoning benchmarks and achieves state of the art streaming pose estimation on ScanNet.
Stable-worldmodel (swm) is presented, an open-source platform for standardized and reproducible world modeling research and evaluation that dramatically reduces research overhead and accelerates trustworthy progress toward reliable world models.
DexHoldem, a real-world system-level benchmark built around Texas Hold'em dexterous manipulation with a ShadowHand, is introduced, which evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting.
PEIRA is introduced, a non-contrastive SSL method with an explicit objective defined through the trace of the optimal linear regressor, which shows that its only stable equilibria are nontrivial global minimizers and recover the same canonical correlation subspaces, with regularization selecting the effective dimension.
Crys-JEPA is introduced, a joint embedding predictive architecture for crystals that learns an energy-aware latent space preserving formation-energy differences that can be reformulated as an embedding-based comparison against accessible training crystals, reducing the reliance on expensive energy evaluation and task-specific external references.
This study advocates semantic latent space as stronger foundation for policy-relevant robotics diffusion world models by comparing six reconstruction and semantic encoders to train world model variants under a fixed protocol on BridgeV2 dataset, and shows effective world model training in high-dimensional representation spaces with and without dimension compression.
This work extends the analysis of Asadi et al. (2018) to MDPs with learned reward models, and derives the optimal sample allocation--the ratio of dynamics samples to reward samples that minimizes a bound on return error under power-law scaling assumptions.
JASTIN is a generalizable, instruction-driven audio evaluation framework that formulates audio assessment as a self-instructed reasoning task that consistently outperforms general MLLMs across speech, sound, music, and out-of-domain evaluation tasks without the need for task-specific retraining.