Portrait of Ruslan Rakhimov

Ruslan Rakhimov

Lead Researcher at T-Tech, Moscow

I work on embodied AI: world models and vision-language-action policies for robots and virtual worlds.

About

I chase the same questions in two domains at once, robots and virtual worlds. The through-line is that hours of video on the internet outnumber recorded hours of robot action by a factor of thousands, and teleoperation scales badly — so I look for recipes that grow a model from cheap video, from its own rollouts under an automatic judge, and from generated tasks that check themselves, rather than from new demonstrations alone. In practice that means world and action models that learn how a world answers an action, and vision-language-action policies that turn a camera image and a plain-language instruction into behaviour. Earlier my research was in 3D computer vision and neural rendering.

I hold a Ph.D. in Computer Science, which I earned at Skoltech within Evgeny Burnaev's Applied AI Center. I received a Master's Degree in Data Science from Skoltech in 2020. Prior to that, I received my Bachelor's in Applied Mathematics and Physics from the Moscow Institute of Physics and Technology (MIPT) in 2018.

News

  • Jul 2026 Released Qantara, a single JEPA checkpoint that serves latent planning, behaviour cloning and inverse dynamics. To appear at an ICML 2026 workshop (DEMO).
  • Jul 2026 Rank-Then-Act has been accepted to the ICML 2026 workshop on RLxF.
  • 2026 Teaching World Models & Vision-Language-Action Models at Central University.
  • 2026 NE-Dreamer has been accepted to an ICLR 2026 workshop.
  • 2026 VL-DAC has been accepted to AAMAS 2026.
  • 2025 GSplatLoc has been accepted to IROS 2025 as an Oral.
  • 2025 Coached the national team of Kyrgyzstan at IOAI.
  • Jan 2025 Joined T-Tech.
  • Dec 2024 Presented the robot Slon with a team at the AI Journey Conference.
  • Sep 2024 Defended PhD dissertation.
  • Sep 2023 Joined Sber Robotics Center.
  • Aug 2023 Our work on reconstruction of 3D meshes of human heads from one view has been accepted to IEEE Access.
  • Jul 2023 Delivered three lectures on Introduction to 3D Computer Vision at the AIRI Summer School.
  • Nov 2022 Became a recipient of The Ilya Segalovich Scientific Award for Young Researchers from Yandex.

Selected Work

All publications
Rank-Then-Act overview 2026

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

ICML Workshop (RLxF), 2026

Yuriy Maksyuta, George Bredis, Ruslan Rakhimov, Daniil Gavrilov

We train a VLM to rank shuffled frames by task progress, then reward a policy by the rank correlation between predicted progress and real time — a bounded, scale-free signal that needs no environment reward and transfers across tasks.

DEF result 2022

DEF: Deep Estimation of Sharp Geometric Features in 3D Shapes

SIGGRAPH, 2022

Albert Matveev, Ruslan Rakhimov, Alexey Artemov, Gleb Bobrovskikh, Vage Egiazarian, Emil Bogomolov, Daniele Panozzo, Denis Zorin, Evgeny Burnaev

Differently from existing data-driven methods for predicting sharp geometric features in sampled 3D shapes, which reduce this problem to feature classification, we propose to regress a scalar field representing the distance from point samples to the closest feature line on local patches.

Multi-sensor 3D dataset sample 2023

Multi-sensor large-scale dataset for multi-view 3D reconstruction

CVPR, 2023

Oleg Voynov, Gleb Bobrovskikh, Pavel Karpyshev, Andrei-Timotei Ardelean, Arseniy Bozhenko, Saveliy Galochkin, Ekaterina Karmanova, Pavel Kopanev, Yaroslav Labutin-Rymsho, Ruslan Rakhimov, Aleksandr Safin, Valerii Serpiva, Alexey Artemov, Evgeny Burnaev, Dzmitry Tsetserukou, Denis Zorin

A new multi-sensor dataset for 3D surface reconstruction that includes registered RGB and depth data from sensors of different resolutions and modalities under a large number of lighting conditions.

DensePose result 2021

Making DensePose fast and light

WACV, 2021

Ruslan Rakhimov, Emil Bogomolov, Alexandr Notchenko, Fung Mao, Alexey Artemov, Denis Zorin, Evgeny Burnaev

We target the problem of redesigning the DensePose R-CNN model's architecture so that the final network retains most of its accuracy but becomes more light-weight and fast.

2021

Latent Video Transformer

VISIGRAPP, 2021

Ruslan Rakhimov, Denis Volkhonskiy, Alexey Artemov, Denis Zorin, Evgeny Burnaev

We predict future video frames in latent space in an autoregressive manner.

Open Source

Patches merged upstream:

  • newton — GPU physics for robotics: making --num-frames terminate a headless GL run, and documenting the pyglet requirement behind it. PRs
  • mjlab — an Isaac Lab API powered by MuJoCo-Warp: setting the GL backend default before MuJoCo is imported. PRs
  • Thea — coding agents for the physical world: running the CLI preflight fixture on the caller's interpreter. PRs

Open and merged, everywhere else: every public pull request I've sent upstream.

Experience

T
Jan 2025 – Present T-Tech Lead Researcher, Embodied AI
S
Sep 2023 – Dec 2024 Sber Robotics Center Lead Research Engineer
Sk
Nov 2019 – Aug 2023 Skoltech Applied AI Center Research Engineer
H
Summer 2019 Huawei Research Intern

Education

Sk
2024 Skoltech Ph.D., Computer Science
Sk
2020 Skoltech M.Sc. with Honors, Data Science
M
2018 Moscow Institute of Physics and Technology B.Sc., Applied Mathematics and Physics