ByteDance

Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

San Jose, CA, USPosted 1 month ago

Job Description

Location

:

San Jose

Team

:

Technology

Employment Type

:

Intern

Job Code

:

A253899

Responsibilities

About the team

The Seed LLM Post Training team is responsible for researching cutting-edge posttrain technologies and providing core posttrain capabilities for unified multimodal large models. The team's goal is to research and explore next-generation advanced technologies such as SFT, RM, RL, and self-learning during the posttrain phase, while significantly optimizing and improving key areas including reasoning, coding, agent, and omni model.

Responsibilities

  • Explore large-scale models and optimize systems.
  • Data construction, instruction tuning, preference alignment, and model optimization.
  • Improving relevant model capabilities, such as reasoning, code, math etc.
  • In-depth research and exploration of future use cases.

Qualifications

Minimum Qualifications

  • Currently pursuing a PhD in Computer Science, AI, or a related field.
  • Research experience in reinforcement learning, sequential decision-making, or agent behavior.
  • First-author publications in accredited ML/AI conferences (e.g., NeurIPS, ICLR, ICML).
  • Solid programming and experimentation skills, including with RL or LLM frameworks.

Preferred Qualifications

  • Experience with LLM agents, tool use, or prompt-based control.
  • Familiarity with environments such as WebArena, ALFWorld, or programmatic reasoning tasks.
  • Understanding of RL techniques such as reward shaping, memory augmentation, or curriculum learning.

As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.

Job Information

【For Pay Transparency】Compensation Description (Hourly) - Campus Intern

The hourly rate range for this

Apply for this role

Keep looking

Similar Remote AI Jobs