STMutants: test quality measured against injected PLC faults
A Structured Text dataset retains 108 mutants from 11 programs; the reported test results are weakest on a stateful sequencer.
Audience and applicability
For verification engineers evaluating generated tests. The PLC benchmark illustrates why an aggregate mutation score needs a separate check of stateful, multi-cycle faults before adapting the method to RTL.
A controlled fault set
Md Humaun Kabir, Md Rakibul Islam and Helen H. Lou introduce STMutants for IEC 61131-3 Structured Text. They create 110 first-order mutants from 11 programs and retain 108 after screening. Each mutant contains one injected change, such as a different comparison, constant, arithmetic operation or initialization. The program set ranges from 38 to 211 lines of code.
The process includes compilation with MATIEC and manual equivalence screening. Compilation establishes syntax and type validity; it does not itself demonstrate a behavioral difference. Screening also depends on observable inputs and outputs, so a fault set and its test assumptions belong together.

What the experiment reports
The abstract calls the second phase kill/survive prediction, but §IV-B gives a more concrete procedure: reuse the generated tests, substitute each mutant into the harness, and execute the original and changed programs using MATIEC. We therefore describe the reported outcomes as tests the authors say they executed, without claiming an independent rerun.
| Test-generating model | All 108 mutants% | Sequence_8% |
|---|---|---|
| GPT-5.2 | 86.1 | 50 |
| Gemini 2.5 | 94.4 | 10 |
| Claude Sonnet 4.5 | 86.1 | 10 |
One-shot generation under the authors' protocol. Sequence_8 has 10 generated mutants and persistent state; the per-program sample is small. These results are not an RTL coverage measurement.
Data sourceThe sequencer result is useful because a high overall score can conceal weak tests of temporal behavior. The paper links failures to deeply nested state-transition conditions. It does not establish that the same scores or operators transfer unchanged to Verilog, VHDL or a particular FPGA design.
A possible RTL experiment
The following is an editorial adaptation proposal. Start with a small owned RTL module whose reference behavior is independently specified. Inject one bounded fault at a time, then ask whether the existing testbench exposes it. Define the observation interval and reset sequence before interpreting a surviving mutant.
- Select operators that match the module's failure modes, such as a changed condition or incorrect state update; do not reuse a PLC taxonomy blindly.
- Run the unchanged design first and verify the expected outputs independently of an LLM-generated test.
- Execute every retained mutant with the same stimulus and record the divergent output and cycle, rather than trusting a textual kill prediction.
- Review surviving mutants for equivalent behavior, unreachable paths and missing temporal stimulus before treating the score as test quality.
Dataset version and missing reproduction pieces
This review uses arXiv v1 from 3 June 2026. The public Figshare artifact checked here is version 3, DOI 10.6084/m9.figshare.31724761.v3, with Mutations.zip. Its metadata declares CC BY 4.0, separately from the article's CC BY 4.0 license. We inspected the archive member list; names did not identify a runnable test harness or result package.
A complete rerun still needs the test vectors, expected-output oracle, execution harness, compiler/runtime revision and model-generation settings. The paper's model names and provider defaults do not freeze an inference service. The dataset is material for a controlled verification study, not a certification of PLC software or FPGA correctness. No LLM or compiler evaluation was run by FPGA.camp.
Sources
Method and test execution protocol
STMutants dataset: pinned version 3