Senior Deep Learning Scientist, Multimodal Agentic RL

nvidia· 2100 NVIDIA USA
Apply Now ↗
📍 US, CA, Santa ClaraFull time

About this role

NVIDIA is widely regarded as one of the technology industry’s most desirable employers. We lead the way in High-Performance Computing, Artificial Intelligence, and Visualization. Our core invention, the GPU, serves as the visual cortex of modern computers and powers our entire product suite. GPU deep learning ignited the modern AI era—the next great computing age—with the GPU acting as the brain for everything from robots and autonomous cars to conversational AI. Today, we are known globally as "the AI computing company." We are looking to grow our teams by bringing in the smartest people in the world. Join us at the forefront of technological advancement.


NVIDIA is hiring Senior Deep Learning Scientists to advance our efforts in streaming and agentic multimodal AI. You will demonstrate foundational expertise in deep learning, reinforcement learning, and applied mathematics to help develop models capable of reasoning, planning, and acting across diverse modalities. This is a chance to define core algorithmic improvements for multimodal foundation models, scaling your ideas through our Nemotron Omni and VoiceChat platforms. You will work on high-impact, high-visibility large language models and multimodal AI products that improve the experience for millions of users. If you are creative and passionate about solving real-world agentic AI challenges, come join our Nemotron LLM team. For more details on Nemotron LLM, check https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/


What you’ll be doing:

  • Apply fundamental and applied research to develop, train, fine-tune, and deploy large language models for agentic systems encompassing audio-visual reasoning, tool usage, and document understanding.
  • Advance post-training and alignment methods including instruction tuning, preference optimization, and RLHF/RLVR/MOPD to improve multimodal agents for complex use cases.
  • Research and develop agentic reasoning and grounded perception capabilities, focusing on planning, tool execution, and long-horizon task completion across digital and physical environments.
  • Lead the collection, development, and benchmarking of multimodal datasets, ensuring high-quality evaluation of model accuracy, safety, and task completion success.

What we need to see:

  • Master’s degree (or equivalent experience) or PhD in Computer Science, AI, or Applied Math with 8+ years of relevant work experience.
  • Excellent programming skills in Python with strong fundamentals in scalable model development and deep learning frameworks like PyTorch.
  • Strong knowledge of ML/DL techniques and modern foundation model architectures, including Transformers and mixture-of-experts models.
  • Foundational understanding of reinforcement learning algorithms and implementation, including MDPs, policies, and reward design.
  • Hands-on experience in post-training multimodal models for omni-modality (audio-visual) reasoning, full-duplex voice chat, and human-AI interaction.
  • Proven ability to manage model development life cycles, including dataset versioning, experiment tracking, and evaluation pipelines.

Ways to stand out from the crowd:

  • Strong record of publications in top-tier AI and machine learning venues such as NeurIPS, ICML, ICLR, or CVPR.
  • Validated experience training and deploying multimodal foundation models using large-scale distributed infrastructure.
  • Experience applying deep reinforcement learning techniques to train multimodal agents in complex simulation or gaming environments.
  • Background in audio/speech AI, especially audio language models or audio generation.
  • Background in building embodied AI systems that integrate multimodal perception with backend action-fulfillment and long-horizon planning.

With highly competitive salaries and a comprehensive benefits package, NVIDIA is considered one of the industry’s most desirable employers. As you plan your future, see what we can offer you and your family at www.nvidiabenefits.com/. If you are a creative and autonomous engineer with a genuine passion for state-of-the-art technology, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 6, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Frequently Asked Questions

Is the salary disclosed for the Senior Deep Learning Scientist, Multimodal Agentic RL position at nvidia?
The salary for this Senior Deep Learning Scientist, Multimodal Agentic RL role at nvidia is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Senior Deep Learning Scientist, Multimodal Agentic RL position at nvidia located?
This Senior Deep Learning Scientist, Multimodal Agentic RL role at nvidia is based in US, CA, Santa Clara. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Senior Deep Learning Scientist, Multimodal Agentic RL role at nvidia full-time or part-time?
This is listed as a Full time position. It is posted as a Senior Deep Learning Scientist, Multimodal Agentic RL role in the 2100 NVIDIA USA department at nvidia.
Which team or department does the Senior Deep Learning Scientist, Multimodal Agentic RL at nvidia belong to?
This Senior Deep Learning Scientist, Multimodal Agentic RL position is part of the 2100 NVIDIA USA department at nvidia. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Senior Deep Learning Scientist, Multimodal Agentic RL position at nvidia?
Click the "Apply Now" button on this page. You will be redirected to nvidia's official application portal hosted on workday where you can submit your application directly.
When was the Senior Deep Learning Scientist, Multimodal Agentic RL job at nvidia posted?
This Senior Deep Learning Scientist, Multimodal Agentic RL position at nvidia was posted on Oct 2, 2026. Apply as soon as possible — early applications are often reviewed first.
Senior Deep Learning Scientist, Multimodal Agentic RL
nvidia
Apply for this role ↗

You'll be redirected to nvidia's official application page on Workday.