Ruslan Rakhimov

I am a Lead Researcher at T-Tech in Moscow, where I work on world models and multimodal VLM agents.

I hold a Ph.D. in Computer Science, which I earned at Skoltech within Evgeny Burnaev's Applied AI Center. I received a Master's Degree in Data Science from Skoltech in 2020. Prior to that, I received my Bachelor's in Applied Mathematics and Physics from the Moscow Institute of Physics and Technology (MIPT) in 2018.

Email  /  Google Scholar  /  X  /  Github  /  LinkedIn

profile photo
News
  • July 2026: Released Qantara, a single JEPA checkpoint that serves latent planning, behaviour cloning and inverse dynamics. To appear at an ICML 2026 workshop (DEMO).
  • 2026: Teaching World Models & Vision-Language-Action Models at Central University.
  • 2026: NE-Dreamer has been accepted to an ICLR 2026 workshop.
  • 2026: VL-DAC has been accepted to AAMAS 2026.
  • 2025: GSplatLoc has been accepted to IROS 2025 as an Oral.
  • 2025: Coached the national team of Kyrgyzstan at IOAI.
  • Jan. 2025: Joined T-Tech.
  • Dec. 2024: Presented the robot Slon with a team at the AI Journey Conference.
  • Sep. 2024: Defended PhD dissertation.
  • Sep. 2023: Joined Sber Robotics Center.
  • Aug. 2023: Our work on reconstruction of 3D meshes of human heads from one view has been accepted to IEEE Access.
  • July 2023: Delivered three lectures on Introduction to 3D Computer Vision at the AIRI Summer School.
  • Nov. 2022: Became a recipient of The Ilya Segalovich Scientific Award for Young Researchers from Yandex.
Work Experience
  • T-Tech: Lead Researcher (January 2025 - Present)
  • Sber Robotics Center: Lead Research Engineer (September 2023 - December 2024)
  • Applied AI Center: Research Engineer (November 2019 - August 2023)
  • Huawei: Research Intern (Summer 2019)
Research

I work on world models and multimodal VLM agents: learning latent models of environment dynamics from pixels and actions, and training vision-language models to act inside them. Earlier my research was in 3D computer vision and neural rendering.

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
Ruslan Rakhimov, George Bredis, Yuriy Maksyuta, Daniil Gavrilov
ICML Workshop (DEMO), 2026
project page / arXiv / code / checkpoints

We train a single JEPA world model whose objective pairs a Brownian-bridge interpolant on the state axis with flow matching on the action axis, so one checkpoint serves latent planning, behaviour-cloning action sampling and inverse dynamics without retraining.

NE-Dreamer: Next Embedding Prediction Makes World Models Stronger
George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov
ICLR Workshop, 2026
project page / arXiv / code

A decoder-free model-based RL agent that predicts next-step encoder embeddings from latent state sequences, matching or exceeding DreamerV3 on the DeepMind Control Suite and improving substantially on DMLab memory and spatial reasoning tasks.

VL-DAC: Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
George Bredis, Stanislav Dereka, Viacheslav Sinii, Ruslan Rakhimov, Daniil Gavrilov
AAMAS, 2026
project page / arXiv / code

A lightweight RL recipe for vision-language agents that applies PPO at the token level for actions while learning a step-level value function. Training in one cheap synthetic environment produces policies that transfer beyond the training simulator.

GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization
Gennady Sidorov, Malik Mohrat, Denis Gridusov, Ruslan Rakhimov, Serkey Kolubin
IROS (Oral), 2025
project page / arXiv / code

We present a novel approach that integrates dense keypoint descriptors into 3D Gaussian Splatting to enhance visual localization, achieving state-of-the-art performance on popular indoor and outdoor benchmarks.

T-3DGS: Removing Transient Objects for 3D Scene Reconstruction
Vadim Pryadilshchikov, Alexander Markin, Artem Komarichev, Ruslan Rakhimov, Peter Wonka, Evgeny Burnaev
Preprint, 2024
project page / arXiv / code

We propose a novel framework that leverages 3D Gaussian Splatting to effectively remove transient objects from input videos, enabling accurate and stable 3D scene reconstruction.

NPBG++: Accelerating Neural Point-Based Graphics
Ruslan Rakhimov, Andrei-Timotei Ardelean, Victor Lempitsky, Evgeny Burnaev
CVPR, 2022
project page / arXiv / code

We take the original NPBG pipeline and make it work without per-scene optimization.

DEF: Deep Estimation of Sharp Geometric Features in 3D Shapes
Albert Matveev, Ruslan Rakhimov, Alexey Artemov, Gleb Bobrovskikh,
Vage Egiazarian, Emil Bogomolov, Daniele Panozzo, Denis Zorin, Evgeny Burnaev
SIGGRAPH, 2022
project page / arXiv code

Differently from existing data-driven methods for predicting sharp geometric features in sampled 3D shapes, which reduce this problem to feature classification, we propose to regress a scalar field representing the distance from point samples to the closest feature line on local patches.

Multi-NeuS: 3D Head Portraits from Single Image with Neural Implicit Functions
Egor Burkov, Ruslan Rakhimov, Aleksandr Safin, Evgeny Burnaev, Victor Lempitsky
IEEE Access
project page / arXiv

We present an approach for the reconstruction of textured 3D meshes of human heads from one or few views.

Multi-sensor large-scale dataset for multi-view 3D reconstruction
Oleg Voynov, Gleb Bobrovskikh, Pavel Karpyshev, Andrei-Timotei Ardelean,
Arseniy Bozhenko, Saveliy Galochkin, Ekaterina Karmanova, Pavel Kopanev,
Yaroslav Labutin-Rymsho, Ruslan Rakhimov, Aleksandr Safin, Valerii Serpiva,
Alexey Artemov, Evgeny Burnaev, Dzmitry Tsetserukou, Denis Zorin
CVPR, 2023
project page / arXiv

A new multi-sensor dataset for 3D surface reconstruction that includes registered RGB and depth data from sensors of different resolutions and modalities under a large number of lighting conditions.

Making DensePose fast and light
Ruslan Rakhimov, Emil Bogomolov, Alexandr Notchenko, Fung Mao, Alexey Artemov,
Denis Zorin, Evgeny Burnaev
WACV, 2021
arXiv / code

We target the problem of redesigning the DensePose R-CNN model's architecture so that the final network retains most of its accuracy but becomes more light-weight and fast.

Latent Video Transformer
Ruslan Rakhimov, Denis Volkhonskiy, Alexey Artemov, Denis Zorin, Evgeny Burnaev
VISIGRAPP, 2021
arXiv / code

We predict future video frames in latent space in an autoregressive manner.

Thanks to Jon Barron for the website template.