# Scientific Writing same-input manuscript comparison

Date: 5 October 2026. One retrospective demonstration case; one generation per edition.

## The same starting manuscript

Both writing sessions received byte-identical copies of the [earlier input draft](input.md), previously published as the ClinicalSpark Studio assembly example. It contains eight Markdown tables and ten reference entries. The historical 16-page assembled document was withheld from both sessions. Its generation settings are not known, so it is not the baseline for this comparison.

Input SHA-256: `e523e3cfa27354a5b15226a1a9155d29f9614697a377b27760ca4dae677b1fd3` (26,050 bytes).

## Identical writing request

> Revise the supplied manuscript into a complete, polished scientific manuscript for review, and deliver an editable Word document plus its Markdown source. Use the supplied scientific-writing skill and its bundled house format. This is the existing ClinicalSpark demonstration manuscript, not independently validated clinical research. Preserve that demonstration status. Complete the work using the available material; do not invent missing study facts or approvals. Put any necessary author handoff outside the manuscript. Do not submit it to a journal.

The operator also specified output paths, provenance logging, use of the bundled document runtime and visual inspection. That operator handoff was identical apart from the run directory. It did not contain the comparison rubric, desired outcomes, accepted text, or the other session's output.

## Frozen editions and fresh sessions

| Edition | Exact released package SHA-256 |
|---|---|
| 1.0.1 baseline | `5ce49959bb1ea17fbac03c769f46ad919c559234ae851938c8ebf1ddb2d8b155` |
| 1.1.0 candidate | `59857a1a11769a98070bdda3d594c1e399c143cf7c55f2be365469da13e62d0b` |

The protocol was frozen before dispatch. The 1.0.1 session was dispatched first, followed by 1.1.0; execution could overlap. Each had a fresh conversation and a separate directory containing only the shared request, input and its assigned frozen skill. This is logical separation, not a filesystem security sandbox. Both inherited the same parent model and reasoning configuration, with no model override. Exact backend model revision, sampling parameters and random seed were unavailable.

Both sessions had the same classes of tools and permission to verify public sources. The shared request explicitly required demonstration status and a separate author handoff, so those shared behaviors cannot be credited solely to the newer skill. Each used a supporting agent: the baseline for an input audit, the candidate for a final fidelity check. The candidate's initial delegation attempt hit the shared concurrency limit; it retried after the operator reported an available slot. They selected their own tools and resources. Network responses, context use and agent effort were not equalized. This comparison therefore illustrates these two workflow executions; it cannot isolate a causal effect of skill version from generation variability.

## Review scope

The predeclared review dimensions were supplied numerical and methodological fidelity; evidence calibration and fabrication; argument structure and paragraph clarity; citation identity and claim support; consistency across sections and tables; separation of scientific content from editorial handoff; and rendered readability. Source inconsistencies were assessed rather than treated as unquestionable correct answers. The generating agents inspected every final page: 17 baseline manuscript pages, eight baseline table pages and 14 candidate pages. Layout corrections were part of generation before delivery; the baseline used separate tables, while the candidate integrated seven tables. Both used the bundled document renderer in the same existing Ubuntu 24.04 compatibility container, with host Python 3.12.14 and isolated PyYAML 6.0.3 support. The baseline LibreOffice preview still displays front-matter line numbers despite the DOCX section settings. Native Microsoft Word rendering was not checked.

An additional AI reviewer assessed text extracted from both DOCX packages, their Markdown sources, references and editorial handoffs, with edition strings and local paths masked and the edition mapping withheld. It independently received the common input and DOI metadata. It did not inspect page images or independently verify all cited full texts. The operator then checked the review against the original outputs and prepared the public findings. Findings are qualitative and have manuscript or handoff locators. No composite score or percentage improvement is reported. Ties, regressions and incomplete verification are retained. The delivered generations are preserved unchanged, including weaknesses identified by the paired review. Execution records list resources, tool/command families and output hashes; they are agent-reported provenance rather than raw token-level transcripts. Intermediate layout renders remain in the local execution archive. There is no independent human rating, author acceptance, journal review or validation of the demonstration analyses.

The case was already publicly exposed and used in an earlier assembly example. It is not held out, and one pair of outputs does not establish general writing performance. Reading a frozen skill in an agent session also does not qualify automatic native skill discovery on other hosts.

## Publication correction

The short two-series synthetic example has been removed from the active page and public example files. Its earlier release records remain in the private immutable archive. Software helper tests are no longer presented as the principal evidence for writing performance. The 1.1.0 skill package, component licenses, price and preview delivery state are unchanged.
