INTERNSHIP DETAILS

Research Intern – Reinforcement Learning for Large Foundation Models

CompanyTencent
LocationSingapore
Work ModeOn Site
PostedAugust 24, 2026
Internship Information
Core Responsibilities
The role involves designing reinforcement learning training recipes for large-scale reasoning models and developing pipelines for autonomous agents. You will also focus on reward modeling and optimization techniques to improve model performance in complex interactive environments.
Internship Type
part time
Company Size
89555
Visa Sponsorship
No
Language
English
Working Hours
40 hours
Apply Now →

You'll be redirected to
the company's application page

About The Company
Tencent is a world-leading internet and technology company that develops innovative products and services to improve the quality of life of people around the world. Founded in 1998 with its headquarters in Shenzhen, China, Tencent's guiding principle is to use technology for good. Our communication and social services connect more than one billion people around the world, helping them to keep in touch with friends and family, access transportation, pay for daily necessities, and even be entertained. Tencent also publishes some of the world's most popular video games and other high-quality digital content, enriching interactive entertainment experiences for people around the globe. Tencent also offers a range of services such as cloud computing, advertising, FinTech, and other enterprise services to support our clients' digital transformation and business growth. Tencent has been listed on the Stock Exchange of Hong Kong since 2004.
About the Role

Business Unit

Technology Engineering Group (TEG) is responsible for supporting the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.

What the Role Entails

Research directions include but are not limited to:

- RL Algorithms for Reasoning Models:Design robust RL training recipes (PPO / GRPO / GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.

- RL for Autonomous Agents: Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies, and self-evolving agent behaviors in complex interactive environments.

-Reward Modeling & Optimization: Develop reward signals and regularization techniques that go beyond outcome-based rewards. Explore token-level reward shaping, entropy-based regularization, and learned reward models that generalize across tasks.

Who We Look For

- Enrolled in a PhD or Master's program in computer science, machine learning or a related field. 

- Solid understanding of RL fundamentals (policy gradients, PPO, GRPO, etc.) and hands-on experience applying them to LLM training.

- Strong programming skills in Python and PyTorch; experience training models on multi-GPU setups, comfortable debugging training instability at scale.

- Ability to independently read, critique, and build on recent research papers.

- Published or submitted first-author papers at top ML/NLP venues (NeurIPS, ICML, ICLR, ACL, AAAI, EMNLP, etc.), or demonstrated equivalent research maturity through preprints and technical reports.

- Familiarity with LLM post-training (RLHF, DPO, GRPO), model merging, or agent frameworks (ReAct, tool-use) is a strong plus.

- Experience with large-scale distributed training (DeepSpeed, FSDP, Megatron) and open-source contributions are welcome.

Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Key Skills
Reinforcement LearningPythonPyTorchLarge Language ModelsPPOGRPODistributed TrainingDeepSpeedFSDPMegatronMachine LearningAgent FrameworksReward ModelingNatural Language Processing
Categories
TechnologyScience & ResearchSoftwareData & AnalyticsEngineering