miHoYo Tsinghua University UCAS HKU

Scaling in GamesContinued Pre-Training for Embodied Agents in Diverse Virtual Worlds

Kuan Zhang1,2,*,‡ Yukun Chen1,3,*,‡ Zhihao Yang1,3,*,‡ Yue Su4 Xiangnan Wu3 Zirong Chen2 Run Luo1,5,‡ Jinkun Hou6 Tao Tan1 Yinhe Zheng1,† Yiming Li2,†

1miHoYo Honkai AI R&D Team 2College AI, Tsinghua University 3University of Chinese Academy of Sciences 4MMLab, The University of Hong Kong 5National University of Singapore 6Peking University

*Equal contribution    †Corresponding author    ‡Work done while interning at miHoYo

Correspondence: yimingli9702@mail.tsinghua.edu.cn, yinhe.zheng@mihoyo.com

One physical world vs. diverse virtual worlds
120training configurations
0.8B–27Bfour Qwen3.5 backbones
50–1,600 hsix gameplay budgets
S1–S5five diversity mixtures
4gameplay benchmarks
2.2×1022training FLOPs

Overview

Does scaling carry over from one physical world to many virtual ones?

Scaling data, diversity and model capacity has improved embodied learning in physical environments, where experience stays grounded in shared physical laws. Virtual worlds break that assumption: every game brings its own visuals, dynamics, controls and objectives. We run a controlled study of continued pre-training (CPT) along three axes, data volume, diversity and model size, and evaluate on gameplay benchmarks that move from in-corpus worlds to unseen distant ones.
The training corpus spans 80 games, and the five mixtures S1–S5 are almost entirely 3D, embodied worlds.

In-corpus · 3DMCU

30 Minecraft tasks across embodied, combat and GUI skills. Minecraft is in the training corpus.

In-corpus · 3DViZDoom

Single-scenario shooting in DOOM, which is also in the training corpus.

Unseen · relatedOmniGameArena

Seven unseen 3D embodied worlds, whose actions have counterparts in training.

Unseen · distantGameWorld

Mostly 2D / 2.5D games: a different view and control regime from the 3D embodied worlds seen in training.

Demo I · One world vs. many

Base model vs. Minecraft-only S1 vs. balanced S3

Same backbone, same 1,600 h budget, different mixtures. S1 trains on Minecraft gameplay only; S3 cuts Minecraft to 27% and fills the rest with other games. Each video is one rollout of the same task, and the label under it reports the result of that rollout: success or failure on MCU, task progress on the other benchmarks. The table below gives the benchmark averages reported in the paper.

Benchmark
Task
Model

Paper results · average time-paused score (%) at 1,600 h ·

Demo II · Frontier comparison

Frontier models pause; the world doesn't

In the time-paused setting the game waits for each decision, isolating reasoning competence. In the real-time setting the game keeps running while the model thinks, so latency becomes part of the score. Frontier models take 4–24 s per step; GameScaling takes 0.13–0.69 s. Each video is one rollout of the same task, and the label under it reports the result of that rollout: success or failure on MCU (a 30 s budget when paused, 3 min in real time), score or task progress on the other benchmarks. The chart below gives the benchmark averages reported in the paper.

Benchmark
Task
Setting

All models

Proprietary Open-source GameScaling (ours) Time-paused

Key findings

What scales, and where it stops

01

Data, diversity and size each help different worlds

More data improves in-corpus play. Diversity improves transfer to related worlds, with limited gains in distant ones. Larger models learn faster but, under single-world training, generalize worse to distant worlds.

02

Interface phase, then world phase

Early training brings broad cross-world gains as the model learns the keyboard–mouse interface. Later gains narrow to in-corpus and related worlds. A three-factor scaling law relates validation loss to data volume, diversity and model size.

03

Small and fast wins in real time

GameScaling is competitive with Minecraft specialists and frontier models in 3D worlds, and scores highest when the world does not pause. Persistent gaps in distant worlds suggest CPT alone is not enough for a generalist agent.

04

More steps help, but only after pre-training

Raising the MCU step cap from 600 to 3,000 lifts GameScaling-9B-S1 from 50.9% to 69.3% and the 0.8B model from 42.3% to 59.3%, with no plateau; the untrained Qwen3.5 backbones stay at 4.0% and 0%.

Benchmark scaling grid
Benchmark performance across data budget, model size and data diversity. Rows: MCU, ViZDoom, OmniGameArena, GameWorld. Columns: 50 to 1,600 hours. Colors: mixtures S1–S5 and the base model; each curve spans the 0.8B–27B backbones.
Loss scaling grid
Loss scaling across data volume, model size and data diversity. Training loss and three-level validation loss for the 0.8B–27B backbones; the last column compares the fitted scaling law with observed loss.

Test-time scaling

More steps help — but only after pre-training

The standard MCU protocol caps every rollout at 600 steps, and all MCU numbers above use that cap. Raising it to 3,000 steps, with no further training, lifts GameScaling-9B-S1 from 50.9% to 69.3% and GameScaling-0.8B-S1 from 42.3% to 59.3%, roughly log-linearly and with no sign of a plateau. The two curves stay about 10 points apart at every budget, so extra steps help a small model about as much as a large one. The untrained Qwen3.5 backbones do not scale this way: Qwen3.5-9B rises only from about 2% to 4.0%, and Qwen3.5-0.8B completes no task at any budget. Continued pre-training raises both the level and the test-time slope.

Test-time scaling on MCU: cumulative success rate against the maximum steps allowed
Test-time scaling on MCU. Cumulative success rate of GameScaling-0.8B-S1 and GameScaling-9B-S1 at 1,600 h (solid) and their untrained Qwen3.5 backbones (dashed) over the 150 rollouts, as a function of the maximum steps allowed, on a logarithmic axis. The standard protocol uses 600 steps.

Citation

BibTeX

@article{zhang2026scaling,
  title   = {Scaling in Games: Continued Pre-Training for Embodied Agents in Diverse Virtual Worlds},
  author  = {Zhang, Kuan and Chen, Yukun and Yang, Zhihao and Su, Yue and Wu, Xiangnan and
             Chen, Zirong and Luo, Run and Hou, Jinkun and Tan, Tao and Zheng, Yinhe and Li, Yiming},
  journal = {arXiv preprint},
  year    = {2026}
}