Shadow Work is an alignment laboratory for complex global systems. It uses large-scale strategic simulation to examine how human and machine reasoning behave under institutional, moral, and geopolitical constraint.
Shadow Work
Shadow Work (GSV-E1, pace-layer E) is an alignment laboratory for complex global systems, registered in the GSV portfolio's games domain. Rather than testing model behavior on isolated prompts, it uses large-scale strategic simulation — a grand-strategy engine spanning 221 countries, thousands of institutions, and interacting pressure types (economic strain, legitimacy crisis, elite fracture, compute supply stress) across a multi-century run — to examine how human and machine reasoning behave under sustained institutional, moral, and geopolitical constraint. The player operates a shadow organization with no win screen: the point is not to optimize toward victory but to observe what a reasoning agent does when handed real, cascading tradeoffs rather than a scripted scenario. The key design decision is that the simulation itself is the instrument — it is the most mature of GSV's simulation projects and serves as the substrate for a dedicated evaluation layer (ShadowBench) built on top of it, meaning the same world model that generates play experience also generates alignment signal, rather than alignment being bolted on as a separate benchmark. This dual-use structure — one simulation, two purposes (play and evaluation) — is part of what GSV frames as "one world, two sapiences, both humbled." Shadow Work connects to a cluster of related GSV projects: it sits alongside elf-revel, agentmud, and duellm in a strategy-simulation lineage, and is grouped with Solwend, elf-revel, after-them, and Unc Rancher in portfolio strategy documents. Current status: it is the most mature simulation in the portfolio, with a Prometheus Crisis demo near-shipping (800+ institutional actors modeled) and an alignment-lab bench (ShadowBench) actively in flight as of mid-2026. It is best understood as infrastructure for a research question — how do human and machine judgment hold up under geopolitical-scale pressure — expressed through playable simulation rather than as a conventional game.