VinMotion

Tuyển dụng

Machine Learning Scientist – Vision-Language-Action (VLA) for Humanoids

  • Hà Nội
  • Full-time

About the job

Overview

  • We're building humanoid robots that can understand instructions, perceive the world, and perform useful physical work in real environments.

  • As a Senior / Staff Machine Learning Engineer, Vision-Language-Action (VLA), you will help build the intelligence stack that connects vision, language, and action. You'll design, train, and deploy multimodal robot policies that allow humanoid robots to understand natural-language commands, interpret complex scenes, and take reliable actions in industrial and logistics settings.

  • This is a hands-on role at the intersection of multimodal model development, robot learning, teleoperation data, simulation, and real-world deployment. A core challenge of the role is learning robust robot behavior from limited real-world teleoperation data, while using simulation and synthetic data to improve generalization and close the sim-to-real gap.

  • You'll work closely with Teleoperation, Simulation, RL & Controls, and Platform teams to turn embodied AI research into robot capabilities that run on physical hardware.

Key Responsibilities

  • Design, train, and improve vision-language-action (VLA) models for humanoid robots using visual observations, language instructions, and robot state

  • Build and fine-tune multimodal robot policies using RGB (and optionally depth), language, proprioception, and task history

  • Improve policy robustness, instruction following, manipulation reliability, and generalization to new tasks and environments

  • Work with the Teleoperation team to improve data quality, dataset curation, logging standards, and episode QA

  • Develop training strategies for limited-data settings, including augmentation, pretraining, imitation learning, and offline RL where appropriate

  • Use simulation and synthetic data to improve robustness and sim-to-real transfer

  • Define evaluation metrics for task success, robustness, and safety, and use ablations / failure analysis to guide model improvements

  • Collaborate with Platform and Controls teams to integrate policies into the on-robot inference stack under latency and memory constraints

  • For Senior-level hires: independently own a workstream from experimentation through deployment iteration

  • For Staff-level hires: help shape technical direction across policy learning, data strategy, evaluation, and sim-to-real roadmap

Required Qualifications:

Core skill

  • Strong background in deep learning for vision, multimodal learning, sequence modeling, or robot learning

  • Hands-on experience training models in PyTorch

  • Strong Python engineering skills, including training pipelines, experiment management, debugging, and reproducibility

  • Experience in at least one of the following:

+ Vision-language / multimodal model training

+ Imitation learning / behavior cloning / offline reinforcement learning

+ Robot learning from demonstration, teleoperation, or real-world interaction data

General

  • Ability to work effectively across ML, robotics, controls, hardware, and platform teams

  • Experience owning technical problems end-to-end, from experimentation through evaluation and iteration

  • BS / MS / PhD in Computer Science, Robotics, Electrical Engineering, or a related field — or equivalent industry experience

Preferred Qualifications

  • Experience training policies for real robots, especially manipulation or articulated systems

  • Experience working with teleoperation or demonstration data

  • Experience with simulation / synthetic-data pipelines for robotics

  • Familiarity with modern VLA and robot-policy approaches such as VLM-backbone policies (Pi0-, GR00T-, or OpenVLA-class), diffusion policy / flow-matching action heads, or ACT-style action chunking

  • Experience deploying models under production constraints such as latency, memory, mixed precision, or on-device / on-robot inference

  • Prior experience in humanoid robotics, embodied AI, robot manipulation, or applied multimodal systems

  • Tech stack includes: PyTorch, ROS 2, Isaac Sim / Isaac Lab, MuJoCo, GPU training infrastructure, and on-robot inference optimization.

Cơ hội việc làm tương tự

AI/Robotics Researcher/Postdoc

  • Fulltime
  • Ha Noi
Hạn nộp
Xem chi tiết

Software Tester (Automation)

  • Full-time
  • Hà Nội
Hạn nộp
Xem chi tiết

Full Stack Engineer, Robotics Simulation

  • Full-time
  • Hà Nội
Hạn nộp
Xem chi tiết