~/satyajit

GTA V in a browser tab: Rockstar's own engine, recompiled to wasm64 and WebGPU

mdjsonmcp

2026-10-06 · 46 min · webgpu · systems · gpu · explainer

A video went round this week of someone playing Grand Theft Auto V in Chrome on an M3 Pro, at playgta5.com. The replies split three ways: it is cloud gaming, it is an Xbox 360 emulator, or it cannot be real. The most common follow-up was how a game of 100 GB or more fits in the "700 MB" the tab downloaded.

It is real, it is none of the first two, and nothing was compressed to 700 MB. I opened the site, downloaded every file it serves, read all of its JavaScript, pulled sample game files apart byte by byte, and ran the game headless on a 16-core server with no GPU. This is what is in the tab, piece by piece, with the code.

What runsRockstar's RAGE engine and GTA V's game code, compiled to WebAssembly (wasm64, shared memory, 75 threads)
Engine binarygame.wasm, 63,201,802 bytes; about 91,000 functions, all still named
GraphicsDirect3D 11 calls recorded into a 32 MB ring as 48 opcodes, replayed on WebGPU by a worker
Shaders4,919 compiled DXBC shaders translated offline to WGSL; 4,808 usable; 72 MB in 18 packs
Game data5,814 files, 20.9 GB, of which 1,528 .rpf archives are 20.3 GB; streamed, never downloaded whole
StorageHTTP Range requests in 4 KB blocks into an append-only store in the Origin Private File System
What a session downloadsa recorded 582 MB boot set, then the blocks the world needs: 743 MB in all, in my run
My run (measured)world on screen at 188 s, 0.3 fps on SwiftShader (a CPU), 2,736 draw calls per frame

One caveat before the engineering, because it shapes everything else: the binary was built from GTA V's own C++ source, and the 20.9 GB the site serves is Rockstar's game data. Neither is licensed for this. I come back to where they almost certainly came from near the end.

What you see

The page opens on GTA V's loading screen: the character art, the logo, sliding in on a tilted 3D plane. This is not a video. It is the game's Scaleform loading movie, LOADINGSCREEN_STARTUP, decompiled from its ActionScript and rebuilt in HTML and CSS: the same 55-degree camera, the same six random orders of 17 screens, a new screen every 14 seconds. The page says so in a comment, down to the font ($Font2 = Chalet London 1960). The art, logo and font are the game's own, extracted by a script the comments call phase3/make_title.py.

Bottom right, a menu offers Story Mode (continue the newest savegame, or the Prologue with ?newgame=1) and Sandbox Mode, which asks for a map: GTA V Map (free roam at a random spot with every weapon) or a second map the page labels "GTA VI Map". That one is the env_test level in the data. You can change your mind until the engine reads its -level argument, about 62% of the way into the load. If you have not chosen by then, the engine waits and the key caps pulse.

The GTA V loading screen in a browser: Michael counting money in front of the Vinewood hills, the GTA V logo bottom left, and at the bottom right the buttons 'Sandbox Mode SPACE', 'Story Mode ENTER' and 'Mounting archives 24%'. A small log line at the bottom reads '[file] Root folder /game/'.
Twenty seconds in, with no mode chosen: the start menu, the progress label driven by the engine's own log, and the newest engine log line at the bottom. The loading screen is the game's Scaleform movie rebuilt in HTML (screenshot of playgta5.com, my headless run).

Under the art, the progress bar is fed by the engine's own log. The port does not modify the engine to report progress. It matches nine regular expressions against the lines RAGE always printed, in prejs.js (the porter's code at the top of game.js):

const stages = [
  [/^Audio rpf =/, 24, "Mounting archives"],
  [/Using Settings/, 38, "Reading settings"],
  [/created D3D11/, 45, "Starting graphics"],
  [/CShaderLib::Init/, 50, "Loading shaders"],
  [/LOAD_ADDITIONAL_TEXT - finished/, 56, "Loading text"],
  [/Starting audio thread/, 62, "Starting audio"],
  [/Loading singleplayer metadata/, 66, "Loading audio metadata"],
  [/^Loading SRL file/, 72, "Starting scripts"],
  [/Removing prologue IPL light groups/, 84, "Loading the world"],
];

The first 20% is the engine download, 76% is the first frame the GPU worker presents, and 72-98% follows the game data that arrives. The title screen drops only when 1.5 s of frames go by with no draw call waiting for a pipeline to compile. "Loaded" is defined by the shader compiler, not the engine.

I ran it with ?mode=sandbox. The timeline, from my log (measured, on a server in a data centre, through a Cloudflare edge in Mumbai, with eight other jobs on the machine):

TimeStageGame data fetched
24.6 sengine downloaded and compiled, main() starts0
39.5 smounting archives(boot set prefetch running)
54.6 sstarting graphics223 MB
69.7 sfirst frame454 MB
176.2 spreparing shaders743 MB
188.1 sworld shown743 MB

Your laptop will compile faster and render far faster. It will not fetch much faster, for a reason the storage section explains.

GTA V gameplay in the browser: a man in a grey suit seen from behind on a tiled plaza in downtown Los Santos, skyscrapers, trees and railings around him, the minimap and health bar in the bottom left.
The world at 188 s on a machine with no GPU at all. Chrome's WebGPU adapter here is SwiftShader, a CPU implementation of Vulkan, which draws this frame's 2,736 draw calls at 0.3 frames per second (screenshot of playgta5.com, my headless run).

What the site actually ships

Before reading any of it, the inventory. Everything is served from one host behind Cloudflare. The engine files live under a build-specific prefix (/b/8b0b5899ed/) and are cached for a year; the game data lives under /data/ with ?v=<manifest version> on every URL, so the same URL always means the same bytes.

FileSizeWhat it is
index.html44 KBthe page: title screen, start menu, input, audio start-up, low-memory profile, engine command line
loader.js9 KBa worker that becomes the engine's "main thread"; starts the GPU and IO workers; streams and compiles the wasm
game.js124 KBEmscripten's glue plus the porter's prejs.js (logging, progress, crash reports) and 13 JS imports; runs in every thread
game.wasm63 MBthe engine: RAGE and GTA V, compiled to wasm64
wgpu_worker.js194 KBthe GPU worker: replays Direct3D 11 commands on WebGPU
io_worker.js41 KBthe IO worker: serves every engine file read from HTTP and a disk store
audio-worklet.js1.5 KBplays the engine's mixer output
shaders/index.json686 KB4,919 shaders: hash to stage, effect, program, pack and offset
shaders/pack0..17.*.bin18 × ~4.2 MBthe translated WGSL, bundled by effect
shaders/pipelines.json359 KB720 render-pipeline recipes to compile ahead (223 KB, 425 recipes, for the low profile)
data/manifest.json393 KBthe game folder: 5,814 paths, sizes and dates
data/bootset.json193 KBthe byte ranges the engine reads while booting: 2,918 files, 582 MB
data/**20.9 GBthe game files themselves

The JavaScript is unminified (except game.js) and commented like a lab notebook. Comments name the porter's own tools (phase3/make_shader_pack.py, pipeline_seed.py, record_bootset.py, a PHP host at phase3/host/index.php), the C++ files on the other side of each interface (wgpu_backend.cpp, httpfs_wasm.cpp, userdata_wasm.cpp, browser_input_wasm.cpp), and measurements with dates. Several are dated 6 October, the day 10,000 visitors overloaded the host. The GPU worker even runs unchanged under Node with Google's Dawn, which is how the porter takes headless screenshots on a Windows dev box (D:/wasm_build/... in a default path).

It is the real engine, not an emulator

An emulator would ship a console's CPU interpreter and a disc image. Cloud gaming would ship a video decoder and a WebRTC connection. This site ships neither. I parsed game.wasm's sections without running it:

SectionSize
code48.2 MB
data7.8 MB
name (function names)6.8 MB
functions defined91,026
imports86

The name section is the giveaway. It was left in, and it lists 91,111 function names. 41,286 of them contain rage::, the namespace of Rockstar's RAGE engine. The rest read like a tour of the game: 1,865 handlers in network_commands, 1,606 in vehicle_commands, 1,392 in ped_commands, 1,162 in hud_commands. Those are the native functions GTA V's mission scripts call. The audio code includes rage::audDecoderPcm, rage::audDecoderAdpcm and rage::audDecoderOpus. A namespace called wasm_null_d3d is the port's own, and comes up in the graphics section.

The strings agree. Assertion messages carry the build machine's source paths, for example E:/P1/GTA5/SRC/DEV_NG\RAGE\BASE\SRC\net/status.h. DEV_NG is the next-gen branch, the one the PC version came from.

So the port was built from C++ source, by a compiler, not by translating machine code. And it was a development build:

The memory is the other tell. The module imports one memory with flags 0x7: it has a maximum, it is shared between threads, and it is 64-bit. Minimum 49,152 pages, maximum 262,144; at 64 KB a page, that is a 3 GB starting heap and a 16 GB ceiling. game.js creates it like this:

var INITIAL_MEMORY = 3221225472;
wasmMemory = new WebAssembly.Memory({ initial: BigInt(INITIAL_MEMORY / 65536), maximum: 262144n,
                                      shared: true, address: "i64" });

Classic 32-bit WebAssembly stops at 4 GB, and a 2013 PC game sized for 8 GB machines does not fit in that. In my run, once the world was up, the engine reported a 2,444 MB malloc arena with 2,405 MB of it in use. Every pointer crosses into JavaScript as a BigInt, and every address is above 3 GB, which caused one of the port's best bugs (see "Graphics").

The game folder

The manifest is a complete listing of the folder the engine sees as /game/. Explore it: the bar is each folder's share of its parent, and the green part is what the recorded boot set reads before the world appears.

20.94 GB · 5,814 files · boot set reads 581.8 MB
From the site's data/manifest.json and data/bootset.json. Grey: the folder's share of its parent. Green: the part the boot read set fetches before the world is up. Below the ten largest entries of a folder, the rest are lumped into one row.

The prose version of the tree:

FolderSizeFilesWhat it holds
x64/levels/gta58.68 GB1,383the GTA V map: _hills 3.95 GB, _citye 1.16 GB, _cityw 0.93 GB, _prologue 0.44 GB, scripts 0.43 GB, vehicles 0.34 GB, interiors, props, navmeshes
x64/levels/env_test5.49 GB259the second map the page calls "GTA VI Map": countryside, co1, sb, sn, nb, dn and 30-odd more blocks
x64/models/cdimages2.50 GB40characters: streamedpeds_players.rpf alone is 532 MB
x64/audio/sfx1.88 GB54sound banks: interactive music 472 MB, cutscene audio 153 MB, 19 radio station banks at 7-52 MB each
x64/anim0.84 GBcutscene and in-game animation
x64/dlcPacks0.58 GB11 packs (see below)
common/shaders0.32 GB1,134compiled effects in three variants: win32_40, win32_40_lq, win32_nvstereo
common/data0.22 GBXML and metadata: AI, timecycles, glass, scripts, paths
common/non_final0.04 GBcutscene lighting tunes (.lightxml) and animation data

By type, 1,528 .rpf archives are 20.30 GB of the 20.94 GB. Then come 1,134 .fxc shader effects (316 MB), 990 .xml files (174 MB, one of them a 145 MB paths.xml), 493 .lightxml, 470 .ymt, 376 .meta, 146 .dat and 65 .ytd texture dictionaries.

What the files are, from samples

I fetched a few sample files with Range requests and read their headers. Nothing was decrypted.

How 100 GB becomes "700 MB"

It does not. Two separate questions are folded into that number.

Why 20.9 GB, not 100 GB? The PC version shipped at 65 GB in 2015 (reported, seven discs) and has grown since with years of GTA Online updates. This folder is a development snapshot from February 2015, and the manifest shows what it leaves out:

I have not diffed it against a retail install, so I cannot say file by file what is smaller and why. What is clear is that this is not the retail game repacked.

Why "700 MB"? Because a session never downloads the 20.9 GB. It downloads the blocks the engine reads, 4 KB at a time. In my run that was 743 MB before the world appeared: 3.5% of the folder (reasoned from the two measured numbers). Most of it is the boot set, which I could break down from bootset.json:

Part of the boot setBytesWhy
.rpf archives433 MB997 of the 1,192 archives it touches are read for under 64 KB each: their headers and tables of contents
.fxc shader effects100 MB349 of the 379 win32_40 effects, fetched whole in compressed batches
x64/levels/gta5258 MBthe map's archive headers and the start-up area
x64/models/cdimages112 MBcharacter model archives
everything elsethe resttexture dictionaries, metadata, AI data, timecycles

(The rows overlap: the level and model bytes are mostly .rpf bytes.) The engine mounts about 1,200 archives at boot and reads only each one's table of contents, "median 3 KB, 4.9 MB for all of them" according to the IO worker's comments. After the boot set, the world needs "about 1.2-1.4 GB the first time" per the page's comments. In my run, at the spawn point I got, it needed far less. Walk or drive somewhere new and more streams in.

Start-up: who runs where

A game engine assumes it owns its threads. It blocks on mutexes, sleeps, spins and waits for the GPU. A browser's main thread may do none of that. So the page does almost nothing, and the engine runs entirely in workers.

  1. The page (index.html) transfers the canvas to an OffscreenCanvas, decides on the low-memory profile, builds the engine's command line, and starts loader.js in a worker. It keeps only input, audio start-up and the title screen.
  2. loader.js creates the GPU worker and the IO worker first, because "Chrome fetches the script of a worker nested in a worker through the parent's event loop", and this worker's event loop is about to be blocked forever. The IO worker starts prefetching the boot set at once, in parallel with the engine download.
  3. loader.js streams and compiles game.wasm while it downloads, then calls importScripts('game.js'). Static constructors and main() run right there, and the game loop blocks this worker for good, which is fine because nothing else lives on it.
  4. Emscripten pre-spawns a pool of pthread workers: min(160, 36 + 4 × cores), or 28 at two cores or fewer. Each runs the same game.js against the same shared memory.

The engine download, from loader.js:

const streaming = async () => {
  const res = await fetchWasm();             // retries 503/500/429 eight times
  const counter = res.clone().body.getReader();  // count bytes on a clone for the progress bar
  ...
  return WebAssembly.instantiateStreaming(res, imports);
};

Streaming compilation means no wait after the download, no second 63 MB copy, and Chrome can cache the compiled code for the next visit.

The command line the page builds, trimmed:

args: ['-rootdir=/game/', '-wgpu', '-nodisplaycalibration', '-noSocialClub', '-pc:nosocialclub',
       '-nonetwork', '-output', '-width=' + w, '-height=' + h, '-forceResolution',
       '-Script_all=warning', '-replay_all=warning', '-fastscriptcore', '-frameLimit=1', ...]

-nonetwork and -noSocialClub remove the online layer. The two log-level flags exist because scripts and replay produce "~87 % of all log lines in play". -frameLimit=1 is a 60 fps cap.

Because the "main" thread is a worker, Emscripten's glue sets can_block = !ENVIRONMENT_IS_WEB, which is true. The engine's main loop can Atomics.wait like any other thread and never returns to an event loop.

Threads it does not start. Of the engine's threads, 75 started in my run and 18 were skipped by the port's thread wrapper: the hang-detect watchdog, Bink video, the TCP and netThrPool network workers, HttpQuery, Bugstar logging, Fwd Thread, Back Thread and six replay-recording threads. ?keep=replay,net,bink,idle brings groups back, as a bit mask the C++ side reads through one of the port's JS imports. The port also has its own hang detector, which prints which thread waits on which semaphore when the engine goes quiet; my logs have several.

The low-memory profile. Devices reporting 4 GB or less (navigator.deviceMemory, which browsers cap at reporting 4 for 4 GB and less) get a lighter game. Every engine thread is a worker of about 20 MB and RAGE sizes its pools from the core count, so the profile tells the engine there are 2 cores, which the page says gives 54 threads instead of 80 to 116. It also quarters texture memory, drops ped and vehicle variety, city density and LOD to minimum, and turns shadows off:

const LOW_ARGS = ['-textureQuality=0', '-pedVariety=0', '-vehicleVariety=0',
  '-shadowQuality=' + (q.get('shadows') === '1' ? 0 : -1), '-reflectionQuality=0',
  '-particleQuality=0', '-grassQuality=0', '-cityDensity=0', '-lodScale=0',
  '-pedLodBias=0', '-vehicleLodBias=0'];

The porter measured it in a Node run: "GPU memory 1.16 GB -> 0.5 GB and streamed assets 1.38 GB -> 0.95 GB". Shadows went because the cascade shadow pass was "a third of all draws (Node, 1366x600: 600-900 of 1,300-2,900 per frame), and a 4-core Chromebook spent ~110 us per draw".

Shared memory: the one rule

The rule that makes the rest work: engine threads never call a browser API that has to yield. fetch, WebGPU and audio are all asynchronous, and a C++ thread blocked in Atomics.wait never returns to an event loop to receive their callbacks. So JavaScript workers with normal event loops do the asynchronous work, and the engine talks to them only through structures in shared memory.

one shared wasm64 memory · 3 GB at start · 16 GB ceilingpick a box
SharedArrayBuffer · AtomicsPageDOM events, title screenEngine threadsmain() + 75 pthreadsGPU workerowns the WebGPU deviceIO workerfetch + OPFS block storeAudioWorklet48 kHz stereo out
in shared memory: 32 MB command ring

The engine still talks Direct3D 11, to a null device that records each call as an opcode plus payload words. The GPU worker drains the ring and replays it on WebGPU with shaders translated offline from DXBC to WGSL. Readbacks and fences wait on a word in the ring.

The four handoffs, with their layouts:

Each worker is the only owner of its browser API, and each handoff is a word in shared memory flipped atomically. The JS imports the port adds to the engine are tiny for that reason. They mostly hand an address to a worker:

function wasm_httpfs_io_start_js(table) {
  if (ENVIRONMENT_IS_PTHREAD || !Module["ioWorker"] || Module["__ioStarted"]) return;
  Module["__ioStarted"] = true;
  Module["ioWorker"].postMessage({ memory: wasmMemory, table: Number(table) });
}

Streaming: a file system made of HTTP requests

The engine's file layer was redirected (platform/file/httpfs_wasm.cpp) so that every read of a file under /game/ becomes a request to the IO worker. First the engine fetches the manifest with a synchronous XMLHttpRequest, so it knows every path and size. Then each read goes through the slot table. Here is one read, end to end, with the IO worker's real code:

one engine read, end to end · engine thread
1. Fill a slot, ring the doorbell, wait

The C++ file layer (httpfs_wasm.cpp) owns one 16-word slot per thread in a table in shared memory. It writes the file id, the 64-bit offset, the destination address in the wasm heap and the length, sets the state word to 1 (requested), bumps the doorbell and blocks in Atomics.wait on its slot. It never touches fetch.

// slot i (16 words at 16 + 16*i): [0] state 0 idle / 1 requested /
//   3 taken by this worker / 2 done, [1] file id, [2..3] offset,
//   [4..5] destination address, [6] length, [7] result (bytes or -1)
io_worker.js:18

In prose: the C++ thread fills its slot (file id, 64-bit offset, destination address, length), sets the state to 1, bumps the doorbell and blocks. The IO worker, parked on the doorbell with Atomics.waitAsync, claims the slot with a compare-and-swap from 1 to 3. It works out which 4 KB blocks of the file are missing, fetches only those with Range requests, writes the bytes into its disk store as they arrive, copies them straight into the engine's buffer in the wasm heap, writes the byte count, flips the state to 2 and calls Atomics.notify. The C++ thread wakes up inside what it thinks was an ordinary read().

The store

The store is two files in the Origin Private File System, opened with synchronous access handles, the one storage API in the browser that reads and writes like a file descriptor:

const root = await navigator.storage.getDirectory();
const dir = await root.getDirectoryHandle('gamedata', { create: true });
const s = await (await dir.getFileHandle('store.bin', { create: true })).createSyncAccessHandle();
...
sh.flush();        // the data first, then the entries that point at it
jh.write(new Uint8Array(buf.buffer), { at: journalEnd });

On the next visit, blocks already stored cost nothing. Sync access handles are exclusive, so a second tab runs without the store, on a 48 MB in-memory cache.

The comments record how the IO worker got here, and the dead ends are the useful part:

  1. Synchronous XHR from each engine thread, behind a service worker that cached blocks. Each response is a fresh ArrayBuffer, freed only when that worker's garbage collector runs. A worker that lives inside wasm or in Atomics.wait almost never runs it, so the bytes piled up as "~1.5 GB of renderer memory", and the service worker added "another ~0.9 GB of its own".
  2. 256 KB blocks, then 16 KB blocks with a growing read-ahead. At boot the engine reads only each archive's table of contents (median 3 KB). Big blocks made "the boot download 25x larger for those files".
  3. Now: 4 KB blocks, streamed into the store with no JavaScript buffers kept, read-ahead only for large sequential reads, and every copy going straight from the store into the shared heap (sh.read(HEAPU8.subarray(...))).

Why so many requests at once

The host answers a Range request at a few hundred KB/s per stream, whatever the visitor's link: "one stream through the host moves ~0.15 MB/s after its round trip of ~0.6 s (a 512 KB read took 2-5 s, measured)". Throughput tracks the number of streams in flight, not bandwidth. The porter's measurements:

fetches in flight → throughput (porter's measurement)
128 in flight → 14.1 MB/s
boot read set · 582 MB41 s
+ world after scripts · 1.2 GB85 s
whole manifest · 20.9 GB24.8 min
Bars on a log scale. Throughputs are the porter's own numbers against the original host; times are those numbers divided into this build's file sizes. A Cloudflare cache hit is faster than this.

At 128 in flight that is 14.1 MB/s, and the boot set takes about 41 s; at 6 in flight it is 3.4 MB/s and nearly three minutes (reported numbers, my arithmetic). That is why your fibre connection does not make the first load much faster, and why every fetch path in the IO worker is built for parallelism:

Failure handling matters at 10,000 visitors a day. Every fetch retries 503, 500 and 429 answers up to 8 times with jittered backoff over about 20 s, because "the engine quits when a data file cannot be read". A file that turns out shorter on the server than the manifest says is shrunk in place, because "the engine retries a failed read forever".

Savegames are separate. When the engine settles a file under /userdata/Documents, the IO worker stores it in IndexedDB (gta5-userdata), not in the OPFS store, which is exclusive to one tab and thrown away when the game data changes. loader.js puts the saves back into the engine's in-memory /userdata before main(), and gives up after 8 s "so a broken IndexedDB cannot keep the game from starting". A failed save is always logged: "a lost save is worth a line".

Graphics: Direct3D 11, replayed on WebGPU

The engine was not rewritten for WebGPU. It still creates a Direct3D 11 device. In this build that device is fake: the function names show wasm_null_d3d::NullDevice, NullContext, NullSwapChain, NullResource and friends implementing D3D11's COM interfaces. Every call becomes an opcode plus payload words in a 32 MB ring in shared memory. wgpu_worker.js is the other end, and its header explains the split:

// Why a separate worker and not calls from the game's render pthread:
//  - WebGPU objects cannot cross threads, but D3D11 resources are created by any engine thread
//    (streaming, main, render). The game threads only emit commands; every WebGPU object lives here.
//  - WebGPU is asynchronous (device creation, mapAsync, onSubmittedWorkDone); this worker has a normal
//    event loop, the game's blocking C++ threads never have to yield to JS. They wait on a futex/fence
//    word in the ring when they need a result (readback, occlusion query).

The ring

Each command is two header words (opcode, payload length) and the payload. There are 48 opcodes, and they are D3D11 almost one for one: CREATE_TEXTURE, CREATE_BUFFER, UPLOAD_BUFFER, UPLOAD_TEXTURE (with the bytes inline in the ring), CREATE_SHADER, CREATE_LAYOUT, CREATE_STATE, CREATE_SRV, CREATE_UAV, SET_SHADER, SET_VERTEX_BUFFERS, SET_INDEX_BUFFER, SET_CBUFFERS, SET_SRVS, SET_SAMPLERS, SET_RENDER_TARGETS, SET_BLEND, SET_DEPTH_STENCIL, SET_RASTER, SET_VIEWPORTS, SET_SCISSORS, DRAW, DISPATCH, CLEAR_RT, CLEAR_DS, CLEAR_UAV, COPY_REGION, COPY_RESOURCE, RESOLVE, GENERATE_MIPS, READBACK, CREATE_QUERY, BEGIN_QUERY, END_QUERY, FENCE, PRESENT, plus WRAP, NOP and a few more. The consumer loop:

async function runCommands(tail, head) {
  while (tail !== head) {
    const base = ringW + (tail >> 2);
    const op = u32[base], n = u32[base + 1];
    if (op === OP.WRAP) { tail = RING_DATA; continue; }
    switch (op) {   // numeric labels: V8 compiles constant Smi cases to a jump table
      ...
    }
    tail += n * 4 + 8;
    if (++sinceYield >= 512) {
      Atomics.store(i32, ringW + 1, tail);   // let the producers reuse the ring space already consumed
      Atomics.notify(i32, ringW + 1);
      if (performance.now() - lastYieldT >= 2) { await yieldTask(); refreshViews(); ... }
    }
  }
  return tail;
}

When the ring is empty the worker sets a "sleeping" flag (header word 5), re-reads the head so a command published in between is not missed, and sleeps on Atomics.waitAsync(head, tail, 100). Producers notify only while that flag is set. When the ring is full, the engine threads block on the tail word, and the C++ side counts that time as "ring full".

Two bugs from this loop are worth retelling:

Resources

Objects live in a plain array indexed by id, because ids come from one counter and are never reused, and "a draw looks up ~15 objects, and an element load is several times cheaper than Map.get". Never-reused ids also make every cache below safe: a key that contains an id can never alias a different object.

Formats are mapped from DXGI numbers to WebGPU formats in a table. Block-compressed textures map directly (bc1 through bc7, which needs the texture-compression-bc feature), depth formats map to depth24plus-stencil8 or depth32float(-stencil8), and typeless formats pick what the game most plausibly views them as, with sRGB view formats added. One format has no WebGPU equivalent: A8_UNORM, an alpha-only texture, which is expanded to RGBA8 on upload because "there is no texture swizzle in WebGPU".

Constant buffers never become GPU buffers. They live in a CPU-side shadow array and are copied per draw (below). Uploads to other buffers keep D3D11's semantics: if a draw already recorded in the open encoder still reads a buffer, the encoder is submitted before the write, so "the earlier draw sees the old data". A Map with NO_OVERWRITE, where the game promises those bytes are unused, skips that flush.

Shaders

Rockstar ships compiled DXBC bytecode, not HLSL source. The port translated every shader offline to WGSL and serves them from shaders/index.json, keyed by the FNV-1a 64 hash of the DXBC, bundled into packs by effect so that a browser makes "~20 requests instead of one or two per shader (~6,000 in a session, each a round trip: minutes against a distant host)". I counted the index:

StageShadersTranslated
pixel3,6673,667
vertex1,1121,112
compute3129
domain710
geometry200
hull180
total4,9194,808

That is 72 MB of WGSL in 18 packs, about 4.2 MB each. The zeros are not failures of effort. WebGPU has no geometry, hull or domain stage. The 109 shaders in those stages are tessellation and the geometry-shader paths of instanced shadows (GS_ShadowInstPassThrough in the cloth effects, for example), so those effects fall back to the engine's other paths or go without.

Here is a real one, fetched from pack0 by its offset in the index: the vertex shader VS_PassthroughComposite of the adaptiveDof effect.

diagnostic(off, derivative_uniformity);
 
var<private> r0 : vec4<f32>;
 
struct cb10_struct {
  tint_symbol : array<vec4<f32>, 10u>,
}
 
@group(0u) @binding(7u) var<uniform> cb7_7 : cb10_struct;
 
var<private> o0 : vec4<f32>;
var<private> o1 : vec4<f32>;
 
fn main_inner(v0 : vec2<f32>) {
  let v = (v0 * cb7_7.tint_symbol[9u].xy);
  r0 = vec4<f32>(v.xy, r0.zw);
  let v_1 = fma(r0.xy, cb7_7.tint_symbol[7u].xy, cb7_7.tint_symbol[7u].zw);
  o0 = vec4<f32>(v_1.xy, o0.zw);
  o0 = vec4<f32>(o0.xy, vec2<f32>(0.0f, 1.0f));
  o1 = fma(v0, cb7_7.tint_symbol[8u].xy, cb7_7.tint_symbol[8u].zw).xyxy;
}
...
@vertex
fn main(@location(0u) v0 : vec2<f32>) -> tint_symbol_1 {
  main_inner(v0);
  return tint_symbol_1(o0, o1);
}

You can read the translation off it. r0, o0 and o1 are DXBC's temporary and output registers, kept as private variables. The constant buffer is register b7 as an array of vec4s, exactly how DXBC addresses constants. The tint_symbol names are what Tint, the WGSL compiler in Chrome's Dawn, emits, so the WGSL was written out by Tint. The bindings follow a fixed scheme: constant buffers by register, samplers at 16 plus the register, textures at 32 plus the register, UAVs at 160 plus, and their hidden counters at 176 plus. The vertex shader's resources go in group 0 and the pixel shader's in group 1.

At load time the worker rewrites that WGSL once more. D3D11 shaders use up to 7 constant buffers per stage, and WebGPU allows 10 dynamic-offset uniform bindings per pipeline layout across all stages. So every stage's constant buffers are merged into one struct at binding 0:

// Every stage's constant buffers ... are therefore merged into ONE struct bound at binding 0 with
// one dynamic offset: `cbN_x` -> `cbpack_v.cbN_x`. The runtime copies the bound D3D constant buffers
// back to back into the uniform ring, in binding order, one 16-byte aligned block per member.
for (const m of members) out = out.replace(new RegExp('\\b' + m.name + '\\b', 'g'), 'cbpack_v.' + m.name);
const struct = 'struct cbpack_t {\n' + members.map((m) => '  ' + m.name + ' : ' + m.type + ',').join('\n')
  + '\n}\n@group(' + group + 'u) @binding(0u) var<uniform> cbpack_v : cbpack_t;\n';

Each draw then copies the bound constant buffers' shadows into a uniform ring. The copies go into a JavaScript mirror of the ring and reach the GPU as one writeBuffer per 8 MB chunk at the next submit, because "a queue.writeBuffer per draw was ~10 % of the worker's time at 1,700 draws per frame". The ring has three frame slots, so two frames can be in flight; with fewer, Chrome's round trip to the GPU process, "not the work, capped the frame rate".

Where the two APIs disagree

The interesting part of a translation layer is never the happy path. These are the D3D11 behaviours the worker emulates, each with a comment explaining what broke:

Pipelines without holes

A render pipeline compiles on first use, and a synchronous compile makes everything submitted after it wait "tens of ms per pipeline, hundreds of new pipelines when a new area streams in: the hitches". So the worker compiles with createRenderPipelineAsync and skips the draws that need a pipeline still compiling. On a slow machine that showed as white holes: "a 4-core Chromebook: ~250 pipelines for the first view of the world took seconds, 28,000 draws skipped in one second".

The fix is to know the pipelines in advance. Every pipeline is also keyed by a recipe: shader hashes, input layout, strides, topology, the raw blend, depth and raster state words, target formats. The site ships a seed of 720 recipes recorded by the porter's Node harness (pipelines.json). Each browser adds the recipes it meets to IndexedDB (gta-pipelines, up to 6,000). While the game loads and the worker is idle, recipes whose shaders have arrived are compiled ahead, a few at a time. My cold run's log line: "title screen held 19.3 s for compiling pipelines ... compiled ahead 719, used 159, compiled when first drawn 37".

Looking a pipeline up has to be cheap, because it happens whenever any pipeline state changes, which is most draws. The key is a tuple of small integers hashed to a 30-bit number:

let h = n * 0x9e3779b1;
for (let i = 0; i < n; i++) { h = Math.imul(h ^ T[i], 0x85ebca6b); h ^= h >>> 13; }
h &= 0x3fffffff;   // a small integer (Smi): a full 32-bit key would be a heap number,
                   // allocated and hashed as a double on every lookup

The string key it replaced "took ~15 % of the worker". The same treatment went to bind groups, cached by an FNV hash of the resource ids in binding order. String keys "were ~40 % of stageBindGroup", and a fast path revalidates the previous group by comparing a few ids instead of rebuilding. The cache is trimmed to 24,000 groups, because "in a long session that exhausted the descriptor heaps (CreateDescriptorHeap failed with E_OUTOFMEMORY, device lost, black screen)": dead groups whose textures streaming had destroyed were held alive by the cache.

Presenting a frame

Render passes begin only when the set of render targets changes, with loadOp: 'load' everywhere (D3D clears are separate passes). At PRESENT, the game's back buffer is copied onto the canvas by a full-screen triangle, the encoder is submitted, and the worker waits for the frame slot it is about to reuse. Readbacks are submitted immediately rather than at present ("deferring the submit to the present made it wait 50-300 ms") and land directly in the engine's staging memory in the heap, followed by an Atomics.notify.

Input, audio and a 1 ms timer

The smaller pieces are each a place where browser behaviour leaks into a game that never expected a browser.

Input. The page maps DOM key codes to the Windows virtual-key codes the engine expects and writes them into the input block:

addEventListener('keydown', (e) => {
  ...
  const vk = VK[e.code];
  if (vk) Atomics.store(keys, vk, 0x80);
  ...
});

It also aliases I/K/J/L/U/O to the numpad keys GTA uses for aircraft pitch, roll and targeting, because most laptops have no numpad. The mouse uses pointer lock with unadjustedMovement, and wheel deltas are converted to notches because "Chrome reports ~100 px per notch on Windows".

Audio. The engine's mixer writes into a ring; an AudioWorklet copies 128 frames per callback and plays silence on underrun. The whole processor:

process(inputs, outputs) {
  const out = outputs[0], left = out[0], right = out[1] || out[0], n = left.length;
  const read = Atomics.load(this.hdr, 1), write = Atomics.load(this.hdr, 0);
  const avail = (write - read) >>> 0, take = Math.min(avail, n);
  for (let i = 0; i < take; i++) {
    const at = ((read + i) & this.mask) * 2;
    left[i] = this.data[at];
    right[i] = this.data[at + 1];
  }
  for (let i = take; i < n; i++) { left[i] = 0; right[i] = 0; }   // underrun: silence
  Atomics.store(this.hdr, 1, (read + take) | 0);
  return true;
}

It starts only after the first click or key press, because browsers do not allow audio before a gesture.

Timers were the strangest bug. Chrome on Windows rounds every wait with a timeout up to the OS timer tick, 15.6 ms, unless the renderer has a timer under 16 ms pending. The engine's sleeps, its 1 ms event waits and its frame limiter all use Atomics.wait with timeouts. Measured by the porter: "a 1 ms wait takes 15.5 ms; with this timer 2.1 ms", and the limiter "held a 60 fps cap to 37 fps". The fix is one line on the page:

setInterval(() => {}, 1);

It does nothing except keep Windows on a fine tick.

Every optimisation, with the porter's numbers

The comments quote a measurement for most changes. Collected in one place:

ChangeWhat it replacedNumber quoted
uniform copies into a JS mirror, one writeBuffer per chunkwriteBuffer per draw~10% of the GPU worker at 1,700 draws per frame
pipeline key as an integer tuple, 30-bit hashstring key per lookup~15% of the GPU worker
bind-group cache by numeric hashstring keys~40% of stageBindGroup
recycling constant-pack entries in place~6 short-lived objects per draw~14% of the GPU worker (garbage collection)
caching GPUBuffer.size in JSa native getter per draw~5% of the GPU worker
MessageChannel yield in the command loopawaiting resolved promises onlya flat 25 fps with the worker half idle
executed-command counter published every 512 commandsan atomic add per commandmeasurable at ~25,000 commands per frame
recipe compile-aheadcompile at first draw28,000 draws skipped in one second on a Chromebook
shader packsone request per shader~20 requests instead of ~6,000
4 KB blocks into OPFS, no JS bufferssync XHR + service worker cache~1.5 GB + ~0.9 GB of renderer memory
table-of-contents-sized blocks256 KB / 16 KB blocksboot download 25x larger for those files
parallel slices for blocking readsone stream per reada 512 KB read took 2-5 s
low fetch priority for speculationequal priorityengine reads waited up to 6.6 s
96 hinted reads in flightthe streamer's one-at-a-time reads~3 reads/s
-fastscriptcorethe debugging script core~8% of the main thread
script and replay logs at warning levelfull logging~87% of all log lines
setInterval(() => {}, 1) on the pageWindows' 15.6 ms timer tick60 fps cap held to 37 fps
low-memory profiledefault settingsGPU memory 1.16 → 0.5 GB

Read together they say where the time goes in a port like this: not in the GPU, but in JavaScript's per-call overheads, the garbage collector and the event loop. Almost every fix replaces an allocation or a string with an integer, or a round trip with a batch.

The control panel in the URL

The port is debugged in production through query parameters, and they are a good index of what the porter had to measure. The useful ones:

ParameterEffect
mode=story, mode=sandbox, map=env_test, newgame=1skip the start menu, start from the Prologue
low=1 / low=0, shadows=1, cores=Nforce or forbid the low-memory profile, keep shadows, fake the core count
res=WxH, scale=0.75, fixed=1, fps=30back-buffer size, render scale, ignore resizes, 30 fps cap
keep=replay,net,bink,idlestart engine thread groups that are normally skipped
slots=N, syncpipelines=1, nopack=1, limits=...frames in flight, synchronous pipelines, one request per shader, a smaller GPU device
dbg=1draw isolation: run only the first N draws of each frame to find the one that paints an artifact
shot=N, shotrt=1screenshots, and a dump of every render target hunting for NaNs
log=1, console=1, verbose=1, mem=1, trace=1, record=1logging, memory reports, a trace of every file read, recording a new boot set
nocache=1, nohints=1no disk store, no read hints

shotrt=1 has a story of its own: it was built for a Chromebook that rendered the world white, where "the lighting buffer was ~95 % NaN, which displays as white". It scans one frame draw by draw and logs the first shader that introduces a NaN.

What it costs

A top-down view of a character in a white suit lying in a pool of blood on a pale pavement next to a reddish tiled area, with a green debug text at the top reading 'C to switch player models'.
Sandbox Mode on env_test, the map the site labels 'GTA VI Map'. The player spawned and died within seconds in my run. The green line at the top is a debug prompt of the development build. This second visit fetched 663 MB in the session, on top of blocks the disk store kept from the run before (screenshot of playgta5.com, my headless run).

Where the source came from

The site does not say, and I cannot prove it. What I can show is what the binary needs: a full C++ source tree for RAGE and GTA V, on Rockstar's own build paths, in a development configuration, with data from a February 2015 build. Rockstar has never published that source.

GTA V's source did leak, twice. Around Christmas 2023, a slice of the material stolen in the September 2022 Rockstar breach went public, and it was reported at the time to contain the full GTA V source, in a form that compiles. In September 2026 a roughly 200 GB archive from the same breach, described as "GTA V source code, debug builds from every platform" plus early GTA 6 development material, began circulating (gtaboom, reported). A development tree with env_test levels, a Bugstar thread and 2014 artist exports is consistent with that material. The "GTA VI Map" label on the menu is the site's, not Rockstar's. I would not read more into it than that env_test is a test level from Rockstar's tree.

The context: ports of a leaked engine

This is not the first port built from that leak, and the previous ones tell you what happens next.

playgta5.com goes further than either. It is not a reverse-engineered engine, and it does not ask you for your own files: it serves a binary compiled from leaked source and 20.9 GB of Rockstar's data to anyone who opens the page. As of 6 October 2026 I could find no press coverage of it, only the viral video. And partway through my last test run, the site's Cloudflare front end started answering this server with "Sorry, you have been blocked". My runs had pulled about 2 GB in many parallel Range requests from a data-centre IP, which is the likely trigger. I did not try to get past it.

Serving all of that to 10,000 people a day is not a gray area. Given what happened to the Switch port a week earlier, I expected the site to be gone soon, and it was: by the evening of 6 October 2026, playgta5.com no longer accepted connections. The Wayback Machine holds a capture of the page from that morning, and it shows exactly how this port was built. The title screen and start menu still render, because they are HTML and CSS. The engine never starts. The archive has no game.wasm, no workers, no manifest and none of the game data (every one of those URLs is a 404 there), and even if it had them, the archived page is served without the COOP/COEP headers it checks for, so it stops at 5% with its own message: "not cross-origin isolated". I have not linked the site or the capture, and nothing here tells you how to get the data.

What I take from it

Strip away the provenance and this is a clean demonstration of something that was not true two years ago. The browser now has every primitive a 2013 AAA engine needs, without rewriting the engine:

What did not translate is equally clear: no geometry or tessellation stages, a per-call cost that makes thousands of draws expensive, garbage collection that never runs in a busy worker, completion callbacks starved by microtasks, timers that round to the OS tick, and a host that has to serve many small streams, not big files. Each of those has a comment in the source explaining the workaround. If you ever port a native engine to the web legitimately, those comments are a better checklist than most documentation.

Elsewhere on the site, WebGPU usually shows up for small things: a 41,321-parameter syntax highlighter, a decision model scored in the page, Gaussian splats as WebP images or TypeGPU's value types. This is the other end of the scale. And for the opposite direction, a modern graphics pipeline reimplemented by reading rather than running it, see OpenDLSS-NR.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "GTA V in a browser tab: Rockstar's own engine, recompiled to wasm64 and WebGPU", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026gta5inthebrowser,
  author = {Satyajit Ghana},
  title  = {GTA V in a browser tab: Rockstar's own engine, recompiled to wasm64 and WebGPU},
  url    = {https://ai.thesatyajit.com/articles/gta5-in-the-browser},
  year   = {2026}
}
share