Journal

STMutants: test quality measured against injected PLC faults

October 9, 2026· 3 min read

A Structured Text dataset retains 108 mutants from 11 programs; the reported test results are weakest on a stateful sequencer.

Audience and applicability

For verification engineers evaluating generated tests. The PLC benchmark illustrates why an aggregate mutation score needs a separate check of stateful, multi-cycle faults before adapting the method to RTL.

A controlled fault set

Md Humaun Kabir, Md Rakibul Islam and Helen H. Lou introduce STMutants for IEC 61131-3 Structured Text. They create 110 first-order mutants from 11 programs and retain 108 after screening. Each mutant contains one injected change, such as a different comparison, constant, arithmetic operation or initialization. The program set ranges from 38 to 211 lines of code.

The process includes compilation with MATIEC and manual equivalence screening. Compilation establishes syntax and type validity; it does not itself demonstrate a behavioral difference. Screening also depends on observable inputs and outputs, so a fault set and its test assumptions belong together.

The authors' Structured Text example with annotations marking individual injected changes.
Mutation example from arXiv v1, CC BY 4.0; rasterized for display. This is PLC code, not an RTL example.Md Humaun Kabir, Md Rakibul Islam, Helen H. Lou — arXiv v1, CC BY 4.0. Rasterized/resized for FPGA.camp.View full-size figure

What the experiment reports

The abstract calls the second phase kill/survive prediction, but §IV-B gives a more concrete procedure: reuse the generated tests, substitute each mutant into the harness, and execute the original and changed programs using MATIEC. We therefore describe the reported outcomes as tests the authors say they executed, without claiming an independent rerun.

Author-reported mutation detection results
Test-generating modelAll 108 mutants%Sequence_8%
GPT-5.286.150
Gemini 2.594.410
Claude Sonnet 4.586.110

One-shot generation under the authors' protocol. Sequence_8 has 10 generated mutants and persistent state; the per-program sample is small. These results are not an RTL coverage measurement.

Data source

The sequencer result is useful because a high overall score can conceal weak tests of temporal behavior. The paper links failures to deeply nested state-transition conditions. It does not establish that the same scores or operators transfer unchanged to Verilog, VHDL or a particular FPGA design.

A possible RTL experiment

The following is an editorial adaptation proposal. Start with a small owned RTL module whose reference behavior is independently specified. Inject one bounded fault at a time, then ask whether the existing testbench exposes it. Define the observation interval and reset sequence before interpreting a surviving mutant.

  1. Select operators that match the module's failure modes, such as a changed condition or incorrect state update; do not reuse a PLC taxonomy blindly.
  2. Run the unchanged design first and verify the expected outputs independently of an LLM-generated test.
  3. Execute every retained mutant with the same stimulus and record the divergent output and cycle, rather than trusting a textual kill prediction.
  4. Review surviving mutants for equivalent behavior, unreachable paths and missing temporal stimulus before treating the score as test quality.

Dataset version and missing reproduction pieces

This review uses arXiv v1 from 3 June 2026. The public Figshare artifact checked here is version 3, DOI 10.6084/m9.figshare.31724761.v3, with Mutations.zip. Its metadata declares CC BY 4.0, separately from the article's CC BY 4.0 license. We inspected the archive member list; names did not identify a runnable test harness or result package.

A complete rerun still needs the test vectors, expected-output oracle, execution harness, compiler/runtime revision and model-generation settings. The paper's model names and provider defaults do not freeze an inference service. The dataset is material for a controlled verification study, not a certification of PLC software or FPGA correctness. No LLM or compiler evaluation was run by FPGA.camp.

Sources

Paper: arXiv v1

Method and test execution protocol

STMutants dataset: pinned version 3

Version3 files and license metadata

CC BY 4.0

More from this section

Article

SourceBack to section