← ATWL supplementary materials
What this page is
The workflow library holds one ATWL representation per paper, but a paper can be represented in several ways. To see how much this matters, we extracted five library workflows again (UTOPIAN, DPVis, EventThread, NodeTrix, SOMFlow) with two LLMs, Claude Sonnet 5 and GPT-5.6-Terra, and compared the results with the library versions and with each other. The experiment is reported in the subsection “Sensitivity to Re-Extraction” of the analysis section of the paper and in the appendix “Details of the Re-Extraction Check”. This page gives the procedure, the results, and the materials needed to inspect and reproduce them.
Main results
- The backbone of a workflow is reproduced. All 15 versions (5 papers × library, Sonnet, Terra) contain a closed cycle, a view that is interpreted, a view that is judged, and a decision that follows; in 12 of the 15 the cycle passes through a decision, a view, and an interpretation. The share of human transforms differs between versions of the same paper by a median of 9 percentage points.
- Detail is not. The number of transforms of a workflow ranges from 9 to 27 across versions. After merging consecutive transforms of the same intent, the re-extractions still have 13 to 27 distinct transitions between intents, the library versions 8 to 16. A re-extraction agrees with the library version in 16–17 of the 25 features, hardly more than two library workflows agree with each other (Table A). The library version of a paper is never the most similar library workflow of its re-extraction (Table B).
- Library-level statistics are stable. Replacing the five library versions by the re-extractions of either model changes the count of any feature by at most 4 of the 54 workflows; the feature frequencies correlate with the original ones at r = 0.98. Only five of 54 workflows were replaced, and the models differ systematically (features per workflow: library versions of the five papers 6.8, Sonnet 7.6, Terra 10.2).
- Reviews. The ten review reports contain 47 critical issues and recommended improvements. By our rough assignment to the phases of the reviewing instructions: 23 structural validation, 18 paper alignment, 5 abstraction level, 1 semantic correctness. GPT-5.6-Terra, reviewing the Sonnet versions, recommended a revision in 5 of 5 cases (15 critical issues); Claude Sonnet 5, reviewing the Terra versions, recommended minor revisions in 4 of 5 cases (2 critical issues). The strictness is confounded with the quality of the versions that were reviewed.
- Reading. Different encodings reflect decisions that the extraction instructions leave to the extractor, for example which part of a tool paper to represent: nine of the ten re-extractions state that they are assumed composite workflows, as the instructions ask for tool papers; none of the 54 library descriptions does. ATWL makes such decisions explicit but does not make them unique, so the representation of a workflow should be checked by its authors or by people who know it well.
| Comparison | Features agreeing (of 25) | Jaccard similarity of present features |
|---|---|---|
| Library version and Sonnet | 17.0 | 0.28 |
| Library version and Terra | 16.0 | 0.30 |
| Sonnet and Terra | 18.0 | 0.43 |
| Two library workflows (1 431 pairs) | 15.6 | 0.22 |
| Paper | Transforms | Features present | Features agreeing | Rank | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| L | S | T | L | S | T | L–S | L–T | S–T | S | T | |
| UTOPIAN | 13 | 10 | 18 | 4 | 4 | 9 | 17 | 18 | 20 | 41 | 9 |
| DPVis | 24 | 21 | 18 | 10 | 10 | 10 | 15 | 15 | 17 | 16 | 9 |
| EventThread | 9 | 13 | 19 | 5 | 7 | 11 | 21 | 15 | 19 | 4 | 32 |
| NodeTrix | 11 | 13 | 14 | 5 | 7 | 9 | 19 | 17 | 17 | 10 | 29 |
| SOMFlow | 27 | 14 | 21 | 10 | 10 | 12 | 13 | 15 | 17 | 27 | 6 |
| Mean (rank: median) | 16.8 | 14.2 | 18.0 | 6.8 | 7.6 | 10.2 | 17.0 | 16.0 | 18.0 | 16 | 9 |
Procedure
- Each of the two models extracted the five workflows one after another in one session, with the extraction instructions, the ATWL definition and three example workflows, as in the library construction.
- The ATWL syntax checker was run on each result and its errors were passed back to the extractor (0–2 corrections per workflow, most often none).
- The other model reviewed the representation against the paper with the reviewing instructions (Sonnet reviewed the Terra versions, Terra the Sonnet versions).
- The extractor revised its representation once. The revised versions contained no syntax errors; they are the files in
atwl/. The first versions, the checker outputs, and the chat logs are not published. - The revised versions were compared with the library versions, which had been extracted earlier. For the five papers, Claude Opus 4.6 took both the extractor and the evaluator role.
The procedure that we recommend for new papers goes further: different models for extraction and review, a syntax check before the first review, review and revision repeated until the reviewer finds no issues, and a final check by an author of the workflow or another person who knows it well. The check reported here used one round of review and revision.
Materials
Instructions and language definition
The same instructions that were used for the library construction:
Explanations of the Sonnet extractor
The notes on the extraction approach, the tables of key extraction decisions, and the summaries of changes after the reviews that Claude Sonnet 5 gave in its session.
ATWL representations
The library versions are copies of the library files that were compared; the Sonnet and Terra files are the versions after one round of review and revision.
| Paper | Library version | Sonnet | Terra |
|---|---|---|---|
| UTOPIAN | library, site page | Sonnet | Terra |
| DPVis | library, site page | Sonnet | Terra |
| EventThread | library, site page | Sonnet | Terra |
| NodeTrix | library, site page | Sonnet | Terra |
| SOMFlow | library, site page | Sonnet | Terra |
Review reports
Reviews of the first syntactically valid versions, written by the other model. Each report is shown as a page and as the original Markdown file. data/review_items.csv lists the 47 critical issues and recommended improvements with our rough assignment to the phases of the reviewing instructions.
| Paper | Version reviewed | Reviewer | Recommendation | Crit. | Impr. | Report |
|---|---|---|---|---|---|---|
| UTOPIAN | Claude Sonnet 5 | GPT-5.6-Terra | Requires revision | 3 | 3 | report, md |
| UTOPIAN | GPT-5.6-Terra | Claude Sonnet 5 | Approve with minor revisions | 0 | 2 | report, md |
| DPVis | Claude Sonnet 5 | GPT-5.6-Terra | Requires revision | 4 | 3 | report, md |
| DPVis | GPT-5.6-Terra | Claude Sonnet 5 | Requires revision | 2 | 2 | report, md |
| EventThread | Claude Sonnet 5 | GPT-5.6-Terra | Requires revision | 2 | 4 | report, md |
| EventThread | GPT-5.6-Terra | Claude Sonnet 5 | Approve with minor revisions | 0 | 2 | report, md |
| NodeTrix | Claude Sonnet 5 | GPT-5.6-Terra | Requires revision | 3 | 4 | report, md |
| NodeTrix | GPT-5.6-Terra | Claude Sonnet 5 | Approve with minor revisions | 0 | 3 | report, md |
| SOMFlow | Claude Sonnet 5 | GPT-5.6-Terra | Requires revision | 3 | 4 | report, md |
| SOMFlow | GPT-5.6-Terra | Claude Sonnet 5 | Approve with minor revisions | 0 | 3 | report, md |
Data and scripts
- features_script_output.csv: the 25 features of the ten re-extractions, as written by
compute_features.py --selected. - library_feature_matrix_2026-10-02.csv: the 25 features of the 54 library workflows (state of 2 October 2026).
- features_recomputed.csv: the 25 features of all 15 versions, recomputed by
scripts/features.py. - versions.csv: size and structure of each version (transforms, artifacts, actors, loops, features).
- pairwise.csv: features agreeing and Jaccard similarity for each pair of versions of a paper.
- rank_of_original.csv: rank of the library version among the 54 library workflows, with the three nearest workflows.
- library_substitution.csv: feature counts of the library when the five library versions are replaced by the re-extractions.
- structure.csv: closed cycles, transitions between intents, and human share of transforms for each version.
- review_items.csv: the 47 review items with their assignment to phases.
- analyse_reextraction.py, features.py, atwlparse.py: the analysis scripts.
Notes
- Reproducing the numbers. Run
python scripts/analyse_reextraction.pyin the folder that containsatwl/anddata/(Python 3, no packages). The features of the re-extractions were computed with the released script; the analysis script recomputes them with an independent implementation that reproduces its output in all 250 cells of the ten re-extractions and the library feature matrix in all 125 cells of the five library versions. The closed cycles and the transitions between intents (structure.csv) agree for all 15 versions with the official packageatwl_tools(longer_patterns.py,graph.intent_graph); see official_check.txt.python compute_features.py --input atwl --selected --out features.csv
- File names. L is the library version, S the Sonnet version, T the Terra version. In the review files,
X_S-review_Tis the Sonnet version reviewed by Terra, andX_T-review_Sthe Terra version reviewed by Sonnet. - Limits. The check covers five papers and two models, with one round of review and revision. The library versions were extracted earlier, by Claude Opus 4.6 in both roles, and had been through the evaluator loop and human supervision; the re-extractions were reviewed by another model and revised once. We did not repeat the same model on the same paper, so the effect of the model and the variation between runs cannot be separated. The assignment of review items to phases is ours and approximate.