2×2×7 Prompting-Condition Experiment (R1)

Full-scale recommendation experiment for the revised ATWL paper

This is the full-scale follow-up (R1) to the 2×2 pilot in experiments, run for the revised version of the paper. Where the pilot compared Formal vs. Prose library representations on two problems with one model, this study crosses 2 design tasks × 2 LLMs (Claude Sonnet 5, GPT-5.6-Terra) × 7 prompting conditions (plus two exploratory extensions). The designs written in ATWL were analysed with 25 features computed from the ATWL text and with the ATWL syntax checker; the prose designs were coded by two LLM coders under a fixed protocol, so that conditions with and without ATWL can be compared.

All session logs, designed workflows (ATWL and prose), prompts, analyses and their data are linked below. See the revised paper for the full write-up; this page indexes the underlying materials.

Materials

Tasks, spec & library

The two design tasks, the ATWL specification and HITL example (linked from the main project pages), and the 54-workflow library in both ATWL and prose form.

Condition protocols & prompts

What each of the 7 conditions supplies and asks for, with the exact prompt text used at every stage.

Workflows

Analysis

Provenance