On October 6, 2026, software developer and tech blogger Simon Willison publicly introduced his open-source project Scrimshaw Jukebox. The experimental platform demonstrates how advanced large language models can be utilized for musical composition far beyond conventional text generation tasks. At the technical core of the system sits Anthropic's Claude Opus 5.5, which is tasked with composing playable soundtracks for video games. Instead of generating raw audio waveforms directly, the model articulates the entire musical arrangement as logical, structured text. Willison published the tool as an open-source repository to highlight new workflows for creative audio production in software development.
Stylistically, Scrimshaw Jukebox deliberately evokes the golden era of classic point-and-click adventure titles. The generated pieces draw inspiration from vintage chiptune soundtracks, made famous by iconic releases such as Monkey Island. These minimalist, polyphonic arrangements rely on distinct melody lines and punchy basslines that historically had to function within very strict sound chip channel limits. For an autoregressive language model, this retro aesthetic provides an ideal testing ground because chiptune note sequences and frequency structures are strictly rule-based and mathematical. As a result, the generated music captures the nostalgic mood of classic gaming adventures with high fidelity.
The core architecture of the tool is anchored by a dedicated plain-text notation format using the .scrim file extension. Whenever Claude Opus 5.5 responds to a composition prompt, it does not output a binary audio file, but rather outputs a clean text score. Within these .scrim files, pitch values, durations, tempos and voice assignments are recorded as transparent, human-readable parameters. This methodology guarantees total determinism and keeps the compositional logic fully auditable. Human developers can inspect the artificial intelligence's compositions in any standard text editor, modify specific notes or rewrite entire measures by hand without friction.
To translate these plain-text scores into audible music, Scrimshaw Jukebox includes a built-in web synthesizer that operates entirely inside the browser. This client-side synthesis engine reads the .scrim data and renders the polyphonic arrangements in real time, eliminating the need for heavy digital audio workstations or specialized local plugins. Users can listen to and iterate on compositions immediately through a lightweight web interface. By separating the symbolic sheet music from the local sound generation, the platform ensures that data payloads remain minuscule across network connections.
This conceptual structure stands in sharp contrast to prevailing audio diffusion models such as Suno. Mainstream generative audio systems typically synthesize opaque audio blobs, outputting a finished waveform that cannot be easily separated into individual tracks or stems. When working with such audio diffusion outputs, developers have virtually no ability to correct a single sour note or adjust the tempo of an isolated instrument. Scrimshaw Jukebox circumvents this black-box limitation entirely. By relying on plain-text score generation, developers preserve granular control over every individual note and voice within the multi-part arrangement.
For independent game developers, this text-first approach provides tangible technical and operational advantages. Because .scrim text files consume almost negligible storage space, indie creators can integrate extensive soundtrack repertoires without inflating game download sizes. Furthermore, the format aligns naturally with procedural audio systems, allowing game engines to assemble, branch or modify musical scores dynamically at runtime. Simon Willison's demonstration highlights an emerging trend across game audio: moving away from monolithic, pre-rendered sound assets toward lightweight, deterministic musical notation powered by general-purpose language models.
