One Source, Many Papers

Publish the source, let each reader’s AI compile the paper, and keep the claims fixed

Every paper I have written was compiled exactly once. I chose the section order, the running example, the notation, and the figure that carries the idea, all for one imagined reader, and the camera-ready deadline froze them. Everyone afterward reads that same build: the ML researcher who has never heard of an isolation level, the engineer deciding whether her team should care, the student trying to reproduce it. Most are not the reader I imagined.

That used to be unavoidable: a second version cost as much as the first, and nobody wrote twelve. It no longer is. So what we publish should change: publish the source, and let each reader’s AI compile the paper for that reader.

The paper is the source; the reader is the target

The published artifact becomes a Markdown file that a reader hands to their own assistant (Claude, ChatGPT, whatever they use) with one sentence: compile this paper for me. The assistant asks how the reader takes in technical material, which programming language they read, what field they come from, and what they want from the paper, then builds one page for them. If it already knows the reader well, it skips the questions and just compiles.

The file has four parts: instructions on how to interview the reader, plan the version, and emit it; a claim ledger listing every number and stated limitation, which every version must keep exactly; the paper itself, the authoritative text; and a fourth I will come back to.

The rules matter more than the prompts. Numbers are fixed. Limitations are never dropped, however short the version. Anything the assistant adds (an analogy, a worked example, a redrawn diagram) is marked as added, never passed off as the authors’ words. Pseudocode may be translated into the reader’s language, line for line and labeled. Plots are never redrawn with invented data. The reader gets a different presentation, not different claims.

Two decisions, two binding times

An earlier post sorted a system’s decisions by one question: when does this stop being open? Ask it of a paper.

A paper makes two kinds of decisions. One is what it claims: the problem, the mechanism, the numbers, the limits. The other is how it presents them: order, example, notation, depth, figures. Today both bind at the same moment, the camera-ready deadline, in the same compiler, the authors, which runs once. Nobody chose to bind them together; there was simply no second compiler.

The two want opposite binding times. Claims should bind early. Others build on them, cite them, and try to break them, which only works if they hold still, and authors can only stake their names on something fixed. Presentation should bind late, because what it depends on (who is reading, what they know, why they opened the paper) does not exist at write time. The author can only guess. The reader’s assistant knows.

Programmers will recognize the move. An ahead-of-time compiler must emit one binary for every machine; a just-in-time compiler waits for the actual machine and workload, and specializes. Nobody thinks the JIT changes what the program means. The same post named the seam that makes this safe: the interface. One level defines it; the level below implements it and can change without renegotiating anything above. In a paper, the ledger is the interface and each rendering an implementation. Fix the interface early, bind the implementation late, and check every implementation against the interface.

So the proposal is not “let AI rewrite papers.” It is narrower: split the paper’s two binding times. An assistant that already knows its reader is the limit of this idea, the latest binding possible, with no interview at all.

One paper, five readers

I tried it on my own paper. Cobra (OSDI’20) checks whether a black-box database actually delivered serializability, and it is dense with graphs, constraint encodings, and a pseudocode figure carrying the core algorithm. I converted the published text into a source file, wrote the ledger (twelve claims, eight limitations), and gave it to five fresh assistants, each seeing only the file and a short reader profile. The source file is public; hand it to your own assistant and get a sixth.

The five papers differed in the ways I hoped:

  • An ML postdoc who prefers diagrams and reads Python got eight thousand words and eleven diagrams, with the hidden version order framed as a latent variable and GPU pruning shown as what it is underneath, repeated matrix multiplication.
  • A systems PhD student who prefers prose, reads C, and wants to judge the claims got the algorithm in C, kernel analogies (lock-ordering checkers, RCU), and reading-group questions on the evaluation.
  • A backend engineer with ten minutes and a Java background got less than half the length, six diagrams, and a section titled “Should your team care?”
  • A programming-languages researcher got the database vocabulary mapped onto memory-model relations and a list of what it would take to rebuild the system.
  • A high-school student with one Python class and almost no systems background. This reader is not the paper’s audience, and the version drops most of the technical content. But it is readable: a key-value store becomes a Python dictionary, a transaction becomes moving ten dollars from Alice to Bob, and the problem becomes a question a teenager can ask: can you trust a database you cannot see inside? The alternative was never a simpler paper. It was an unreadable one.

What did not change is the point. Every number in all five versions traces to the source. All five say the 10× improvement is verification cost, not throughput, the sentence I would most expect a summarizer to get wrong. All five keep the limitations, including the plainest: we did not find a serializability violation in the wild.

A paper that keeps reading after it is published

The fourth part of the file is the one I most wanted.

A published paper is frozen at the literature it read; a reader opening Cobra today cannot learn from it what happened next. So the source carries a dated section labeled after publication: not part of the paper, not written by the authors. For Cobra it records 72 citations on Semantic Scholar as of September 2026 and the work that builds on it directly: snapshot-isolation checkers reusing its encoding, weak-isolation checkers comparing against it, a version adding SQL predicates with a reported 60–100× speedup, and others, plus one paragraph on where the field stands. An assistant that can browse may update it, as long as every new item carries a link and a date and stays visibly separate from the snapshot.

Research already runs on stale reads. This lets a paper refresh its own.

What the author still owns

The obvious objection: if every reader sees a different paper, what did the authors publish, and what are they accountable for?

The source, and they stake their names on the ledger: the same answer I gave about what stays valuable when generation gets cheap, the claim you would pay for if it were wrong. Prose, order, examples, and diagrams are presentation, now cheap to regenerate. Claims are not. Citation and review follow: cite the source and its version, as you would a commit; review the source and its ledger. A rendering is correct when the source could have produced it.

I will admit the cost. A version I never read can still go wrong in ways the rules do not anticipate, such as an analogy that imports an assumption the paper never made. The ledger catches altered numbers and dropped caveats, not misleading metaphors. My bet is that this risk is smaller than the one we already take: a single version most of its readers do not fully understand.

An intermediate stage

This is not the end state. A compiled paper is still a document you read from the top. The version I want answers back: change a parameter and watch the figure change, ask why a step holds and get the proof. That is a later post.

To conclude, I argue for a publication pattern that binds what a paper claims once, early, by the people who stand behind it, and binds how it reads late, once per reader. One source, many papers: not many truths, but one set of claims and as many ways into them as there are readers.