HUMANOID ROBOT

Physical AI
Humanoid robot embodied intelligence research
Physical AI · Multimodal learning · Whole body

Humanoid Robot

We combine motion understanding, language-conditioned skills, locomotion, and manipulation into embodied systems that can learn from people and operate around them.

Human motionVLAWhole-body control
99.8%action-recognition accuracy
16fine-grained actions
7,680IMU samples
2023–26foundations to Physical AI

From understanding human movement to learning useful whole-body robot skills

Humanoid intelligence is more than a walking controller attached to two dexterous hands. The robot must align perception with task intent, coordinate contact across its whole body, and reuse knowledge when objects, language, and environments change.

RV Lab approaches this challenge through Physical AI: representations and policies grounded in sensing, action, and interaction. The program builds on verified work in fine-grained human-action recognition, cooperative control, safe reinforcement learning, simulation, and multimodal demonstration data.

Understand people, learn tasks, coordinate the body

Each pillar supplies a capability required by general-purpose humanoid systems.

01 / UNDERSTAND

Human-motion intelligence

Represent rapid full-body movement using skeletal, inertial, and visual signals across subjects and viewpoints.

02 / FOLLOW INTENT

Vision-language-action skills

Connect observations and natural-language goals to reusable manipulation and transport behavior.

03 / EMBODY

Whole-body autonomy

Integrate balance, reachability, locomotion, and manipulation under contact and stability constraints.

Perception becomes grounded through action

The program connects data, task meaning, learned skills, and safe physical execution.

01 · SENSE

Multimodal observation

Vision, IMU, joint state, contact, language, and environmental context.

02 · REPRESENT

Embodied state

Motion, objects, task intent, and affordances share a learnable representation.

03 · LEARN

Reusable skills

Demonstrations, simulation, and VLA training connect goals to action sequences.

04 · CONTROL

Whole-body execution

Balance, reach, manipulation, and navigation are coordinated under safety limits.

Encoding fast human motion for machine understanding

Readable, native-ratio figures show how wearable signals are transformed into visual representations for classification.

IMU action recognition pipeline for Taekwondo unit actions
Action recognition · Sensors 2024

Turning motion profiles into action images

The study converts multichannel wearable-IMU signals into time-warped action images. A CNN can then distinguish subtle high-speed Taekwondo unit actions that differ in timing and motion profile.

  • Forty expert participants and sixteen unit actions
  • 7,680 samples in the reported dataset
  • 0.998 accuracy with 0.982 F1 score
Open paper and full-size figures →

Figure 1 from the linked open-access article (CC BY 4.0).

Normal and time-warped action image comparison
Representation · Temporal alignment

Preserving the signature of rapid movement

Time warping aligns motion phases while retaining class-specific patterns. This provides a foundation for recognizing human demonstrations before translating them into robot-compatible skills.

  • Wearable signals become compact image-like representations
  • Temporal variation is normalized without discarding motion structure
  • Fine-grained recognition supports demonstration understanding
Inspect the full method →

Figure 7 from the linked open-access article (CC BY 4.0).

From motion data to humanoid task learning

Current work assembles the surrounding data and control stack needed for credible Physical AI experiments.

Human motion sensing and smart fitness application
Human behavior

Motion understanding

Inertial, skeletal, and visual representations of skilled movement.

Humanoid robots handling an automotive component
Whole body

Industrial Physical AI

Locomotion and manipulation grounded in real production tasks.

Bimanual robots collecting task demonstrations
Training data

VLA demonstrations

Multimodal records connecting task language, visual state, and robot action.

Perception, learning, and embodiment

Completed publications are explicitly presented as enabling foundations; the humanoid program itself remains an active direction.

Sensors · 2024

Time-Warped Motion Profiles for Action Recognition

Wearable IMU data become action images that preserve subtle, high-speed motion characteristics.

Read publication →
Sensors · 2023

Viewpoint-Agnostic Taekwondo Action Recognition

Synthetic 2D skeletons projected from 3D motion data improve recognition across camera viewpoints.

Read publication →
KSME · 2026

Fine-Tuning a GR00T VLA Foundation Model

Current work examines small-data limitations using paired simulation and real-world robot datasets.

View proceedings →

Paper figures above come from the linked open-access article under CC BY 4.0. Claims about completed results are limited to the cited action-recognition studies; humanoid-system integration is labeled as ongoing work.

Humanoids integrate the full lab stack

Help build robots that learn in the physical world.

We welcome research in human-motion intelligence, multimodal learning, VLA, simulation, and whole-body robot control.