I chase the same questions in two domains at once, robots and virtual worlds. The
through-line is that hours of video on the internet outnumber recorded hours of robot action
by a factor of thousands, and teleoperation scales badly — so I look for recipes that grow a
model from cheap video, from its own rollouts under an automatic judge, and from generated
tasks that check themselves, rather than from new demonstrations alone. In practice that
means world and action models that learn how a world answers an action, and
vision-language-action policies that turn a camera image and a plain-language instruction
into behaviour. Earlier my research was in 3D computer vision and neural rendering.
I hold a Ph.D. in Computer Science, which I earned at Skoltech within
Evgeny Burnaev's
Applied AI Center.
I received a Master's Degree in Data Science from Skoltech in 2020. Prior to that, I received
my Bachelor's in Applied Mathematics and Physics from the
Moscow Institute of
Physics and Technology (MIPT) in 2018.
News
Jul 2026Released Qantara, a single JEPA checkpoint that serves latent
planning, behaviour cloning and inverse dynamics. To appear at an ICML 2026 workshop
(DEMO).
Jul 2026Rank-Then-Act has been accepted to the ICML 2026 workshop on
RLxF.
2026Teaching World Models & Vision-Language-Action Models at Central
University.
2026NE-Dreamer has been accepted to an ICLR 2026
workshop.
Ruslan Rakhimov, George Bredis, Yuriy Maksyuta,
Daniil Gavrilov
We train a single JEPA world model whose objective pairs a
Brownian-bridge interpolant on the state axis with flow matching on the action axis, so
one checkpoint serves latent planning, behaviour-cloning action sampling and inverse
dynamics without retraining.
Yuriy Maksyuta, George Bredis, Ruslan Rakhimov,
Daniil Gavrilov
We train a VLM to rank shuffled frames by task progress, then
reward a policy by the rank correlation between predicted progress and real time — a
bounded, scale-free signal that needs no environment reward and transfers across tasks.
George Bredis, Nikita Balagansky, Daniil Gavrilov,
Ruslan Rakhimov
A decoder-free model-based RL agent that predicts next-step
encoder embeddings from latent state sequences, matching or exceeding DreamerV3 on the
DeepMind Control Suite and improving substantially on DMLab memory and spatial reasoning
tasks.
George Bredis, Stanislav Dereka, Viacheslav Sinii,
Ruslan Rakhimov, Daniil Gavrilov
A lightweight RL recipe for vision-language agents that applies
PPO at the token level for actions while learning a step-level value function. Training
in one cheap synthetic environment produces policies that transfer beyond the training
simulator.
Gennady Sidorov, Malik Mohrat, Denis Gridusov,
Ruslan Rakhimov, Sergey Kolyubin
We present a novel approach that integrates dense keypoint
descriptors into 3D Gaussian Splatting to enhance visual localization, achieving
state-of-the-art performance on popular indoor and outdoor benchmarks.
Alexander Markin, Vadim Pryadilshchikov, Artem Komarichev,
Ruslan Rakhimov, Peter Wonka, Evgeny Burnaev
We propose a novel framework that leverages 3D Gaussian
Splatting to effectively remove transient objects from input videos, enabling accurate
and stable 3D scene reconstruction.
Albert Matveev, Ruslan Rakhimov, Alexey Artemov,
Gleb Bobrovskikh, Vage Egiazarian, Emil Bogomolov, Daniele Panozzo, Denis Zorin, Evgeny
Burnaev
Differently from existing data-driven methods for predicting
sharp geometric features in sampled 3D shapes, which reduce this problem to feature
classification, we propose to regress a scalar field representing the distance from
point samples to the closest feature line on local patches.
Oleg Voynov, Gleb Bobrovskikh, Pavel Karpyshev, Andrei-Timotei
Ardelean, Arseniy Bozhenko, Saveliy Galochkin, Ekaterina Karmanova, Pavel Kopanev,
Yaroslav Labutin-Rymsho, Ruslan Rakhimov, Aleksandr Safin, Valerii
Serpiva, Alexey Artemov, Evgeny Burnaev, Dzmitry Tsetserukou, Denis Zorin
A new multi-sensor dataset for 3D surface reconstruction that
includes registered RGB and depth data from sensors of different resolutions and
modalities under a large number of lighting conditions.
Ruslan Rakhimov, Emil Bogomolov, Alexandr
Notchenko, Fung Mao, Alexey Artemov, Denis Zorin, Evgeny Burnaev
We target the problem of redesigning the DensePose R-CNN model's
architecture so that the final network retains most of its accuracy but becomes more
light-weight and fast.