This is the full-scale follow-up (R1) to the 2×2 pilot in experiments, run for the revised version of the paper. Where the pilot compared Formal vs. Prose library representations on two problems with one model, this study crosses 2 design tasks × 2 LLMs (Claude Sonnet 5, GPT-5.6-Terra) × 7 prompting conditions (plus two exploratory extensions). The designs written in ATWL were analysed with 25 features computed from the ATWL text and with the ATWL syntax checker; the prose designs were coded by two LLM coders under a fixed protocol, so that conditions with and without ATWL can be compared.
All session logs, designed workflows (ATWL and prose), prompts, analyses and their data are linked below. See the revised paper for the full write-up; this page indexes the underlying materials.
Materials
Tasks, spec & library
The two design tasks, the ATWL specification and HITL example (linked from the main project pages), and the 54-workflow library in both ATWL and prose form.
Condition protocols & prompts
What each of the 7 conditions supplies and asks for, with the exact prompt text used at every stage.
Workflows
-
Designed workflows, all conditions × tasks × models
ATWL representations (C2, C3, C6, C6++, C7) rendered with syntax highlighting and phase navigation, and the prose recommendations (C1–C4) as PDF transcripts.
-
Library-paper & method comparison (C3 vs. C4, C4 vs. C5)
Which library papers each C3/C4 design drew on and what was taken from them, and which methods named across C4 and C5 were included in each — by task and model, with links to each paper's library page.
Analysis
-
ATWL-side analysis
The 25 formal features, the ATWL syntax checker and size descriptors for the 20 ATWL designs; comparison with the library and between conditions, models and tasks.
-
Coding of the prose designs
Protocol, inputs and outputs of the two LLM coders for the 8 prose designs (C1, C4), with data and script; reliability of the coding.
-
Comparisons enabled by the coding
C4 against C1, model and task, the 2×2 across prose and ATWL designs, the lineage from C4 to C6, C6++ and C7, and the explicitness of prose and ATWL.
Provenance
-
Full chat archive (38 sessions)
Every design-task session (and six exploratory feature-extraction sessions), rendered as a readable transcript alongside the raw JSON log, with branch/fork relationships noted.