Researcher, World Models

4 days ago

Santa Rosa, CA, United States Auxo Talent Full-time

Researcher, World Models (Humanoid Robotics)

Location: Bay Area

Salary: $120-180k + Equity


About the Role

We're building the world models that let a humanoid robot perceive, predict and act in the real world. We're looking for a Researcher to help advance that core capability, working at the intersection of self-supervised representation learning, predictive architectures and embodied control, in close collaboration with our platform, firmware and hardware teams.

This role suits someone early in their research career, roughly a year or so in, who's ready for genuine ownership rather than a narrow, tightly scoped lane.


What You'll Do

  • Design, train and rigorously evaluate world models that let the robot predict the consequences of actions across visual, proprioceptive and force/torque modalities
  • Advance our self-supervised learning stack for visual and sensor representations, building on and extending the JEPA family (V-JEPA, I-JEPA and related predictive-embedding approaches)
  • Prototype and benchmark generative and predictive architectures (diffusion, DiT, flow matching, VAEs) against JEPA-style objectives for embodied prediction and planning
  • Own the data pipeline for your experiments end to end, including curation, tooling and scaling, without depending on a separate data-engineering team to move
  • Integrate what you build with our platform, firmware and software teams so your research reaches the robot, not just the paper
  • Contribute to sim-to-real transfer, inverse dynamics and multi-modal sensor fusion, and publish or open-source work where it strengthens the field and the team


What We're Looking For

  • A proven modelling track record: you've trained models and can show solid, honest evaluations, not just training curves
  • JEPA fluency: you understand the joint-embedding predictive approach and can reason about where it fits versus alternatives
  • Breadth across approaches, including familiarity with VLA (vision-language-action) models and a view on their trade-offs
  • Depth in at least one sensory modality: vision, audio, natural language or similar
  • Strong data abilities: you get things done without depending on a whole data-engineering team
  • Solid engineering: you can implement, integrate and ship what you build alongside platform, firmware and software teams
  • A humanoid robotics background, ideally hands-on, and roughly a year into your research career


Nice to Have

  • Publications at NeurIPS, ICML, ICLR, CoRL or RSS (or arXiv work with comparable traction)
  • A PhD or equivalent research experience in ML, robotics or computer vision; not required with a strong portfolio
  • Demonstrated hardware or robotics interest or hands-on experience
  • Strong communication: technical blogs, talks or clear written research