Skip to content

feat(lang): add IEC 61131-3 Structured Text (.st) - #2397

Open
SarletteD wants to merge 3 commits into
DeusData:mainfrom
SarletteD:feature/upstream-st-1a-language
Open

SarletteD wants to merge 3 commits into
DeusData:mainfrom
SarletteD:feature/upstream-st-1a-language

Conversation

@SarletteD

Copy link
Copy Markdown

Refs #1810. Stacked on #2396 (prerequisite style(mcp): compare nullable strings through one helper,
branch fix/cppcheck-nullable-strcmp); until the prerequisite merges, this PR also shows its commit.
This is the first of four Structured Text PRs (1a language, 1b dialect normalizer, 1c export layout,
1d real-world dialect).

What does this PR do?

Adds IEC 61131-3 Structured Text as CBM_LANG_ST on .st, with the vendored grammar
HeytalePazguato/tree-sitter-iec61131-3-st
@ 00e24f50f8de (MIT, (c) 2026 Jorge Centeno), and the ST-specific extraction the graph needs:

  • lang_specs.c: node types for functions/methods (method_signature for INTERFACE members), the
    class-like POUs (FUNCTION_BLOCK, PROGRAM, INTERFACE, type_definition), properties, calls
    (call_expression, invocation_statement), USING imports, branches, variables, assignments.
  • extract_defs.c: every base of EXTENDS/IMPLEMENTS is collected by st_base_classes (the grammar
    repeats the field per base, the generic probe kept one); a small class_base_names dispatches ST there
    and leaves extract_base_classes unchanged; find_class_body answers ST before the generic body probe, which
    would otherwise take a trailing body statement for the member container.
  • extract_usages.c: declared VAR names bind (plural names field), a declaration's initial_value
    stays a read; that check and the existing type-field check now share one helper,
    is_non_binding_decl_field.

Credit: @ysfsmet for the ST groundwork and the discussion in #1810.

In scope / out of scope

  • In: the grammar, language registration, count surfaces, definition/base/usage extraction, tests.
  • Out (following PRs): vendor dialect normalization (1b), the export layout with members after
    END_FUNCTION_BLOCK (1c), the remaining real-world dialect (1d). Without them most real exports parse
    only partially (see the table: 86 % of the proof project's files at this stage).

Vendoring checklist

  • internal/cbm/vendored/grammars/st/LICENSE (MIT, upstream text)
  • THIRD_PARTY.md entry
  • SBOM: generated dynamically from the vendored tree, nothing to edit
  • internal/cbm/vendored/grammars/MANIFEST.md: paragraph + table row + recounted summary/ABI lines;
    "Vendored unmodified; vendor dialect is normalized in internal/cbm/st_dialect.c before parsing."
  • scripts/vendored-checksums.txt regenerated; make -f Makefile.cbm security passes
  • Count surfaces 162 -> 163: README, docs/index.html, docs/llms.txt, pkg/npm/README.md, chocolatey
    nuspec, scripts/package-release.sh, server.json. Four places that still read 158 at the base
    (README.md:229, docs/index.html:57 and :451, docs/llms.txt:8) are corrected too; the language-count
    contract regex does not cover them.
  • Same wiring as the Chialisp precedent: LANG_NAME_TABLE in userconfig.c, a
    CBM_GRAMMAR_CASES fixture, a LABEL_GOLDENS row, lang_ext_st, call-node manifest and language
    registry repro rows (the manifest assert message now states the real count, 228).

Commits

  1. feat(lang): add IEC 61131-3 Structured Text (.st)
  2. feat(st): extract definitions, bases and usages from .st

Size

Production +141 / -8 (extract_defs.c +63 / -1, extract_usages.c +29 / -7, lang_specs.c +39,
registration +10), tests +249 / -11, docs and metadata +64 / -30 (including MANIFEST.md and
vendored-checksums.txt), vendored grammar +97,336 (generated parser.c and friends, excluded from review
size). Per commit: language +49 prod / +67 -11 test; extraction +92 / -8 prod, +182 test.

Checks

  • Focused suites (extraction grammar_regression grammar_labels language lang_contract userconfig): all pass (602 PASS lines) except the known environment failure test_extraction.c: sp is NULL; the first commit alone builds and passes the same suites
  • make test-repro: the same 11 suites / 15 tests fail as on main in this environment (issue480, 523, 546, ... - HTTP/network dependent); the ST repro rows pass
  • scripts/lint.sh --ci: rc 0
  • lint-tidy-diff (clang-tidy 21.1.6, diff against the prerequisite): one finding, in upstream code the
    change has to touch: is_binding_occurrence is over the cognitive-complexity threshold on main (33).
    This PR lowers it to 32 (the type-field check and the ST initial_value check now share
    is_non_binding_decl_field), but the diff-scoped run still reports the function because a changed line
    sits inside it. extract_base_classes is left byte-identical to main.

Proof on a real project

Stage partial .st files nodes edges index time, median (3 runs)
upstream main @ 64c23fa .st not indexed 297 300 7.67 s (7.67 / 7.67 / 7.40)
1a language (4c7bc08) 822 / 956 (86.0 %) 3430 8320 8.65 s (8.43 / 8.93 / 8.65)
1b normalizer (bb6c46e) 263 / 956 (27.5 %) 3431 8153 7.94 s (7.94 / 7.94 / 7.93)
1c export layout (9ec9502) 156 / 956 (16.3 %) 5045 17550 7.94 s (7.94 / 7.92 / 8.18)
1d real-world (49e8e9c) 6 / 956 (0.6 %) 5108 17575 7.93 s (7.94 / 7.93 / 7.92)

Machine: Linux 6.18.33.2-microsoft-standard-WSL2 x86_64 (WSL2 on Windows 11), 13th Gen Intel Core i9-13950HX, 32 logical CPUs, 15 GiB RAM visible to WSL. Release CLI (make -f Makefile.cbm cbm, -O2), cli index_repository on SRCI @ 67fd773 (956 .st files), fresh CBM_CACHE_DIR and CBM_RUNTIME_DIR per run, median of 3 wall-clock runs, nothing else running. Wall time includes process start and the non-ST files; it stays flat because the ST files are small.

Edges by type, 1b -> 1c: CALLS 104 -> 4007, DEFINES_METHOD 2 -> 1606, USAGE 4512 -> 5803, WRITES 0 -> 460, OVERRIDE 0 -> 401. 1c -> 1d: DEFINES 3740 -> 3803 (types that parse now), USAGE 5803 -> 5765 (struct field names no longer read as usages). 1a -> 1b: edges 8320 -> 8153, of which USAGE 4681 -> 4512.

Known limitations (this stage)

  • .st is also used by Pharo/Smalltalk Tonel (*.class.st) and StringTemplate; such files would be
    parsed as ST and yield little or nothing.
  • Real vendor exports mostly parse partially until 1b-1d land (86 % partial on the proof project here).

Written with Claude Code (Anthropic) on behalf of @SarletteD, who reviewed and signs off every commit.

SarletteD and others added 3 commits September 27, 2026 01:40
cppcheck 2.20 (the version the lint job pins) reports
knownConditionTrueFalse on two comparators in src/mcp/mcp.c whose
tie-breakers compare one strcmp result after another, each call with
inline "x ? x : \"\"" fallbacks. Neither function differs between these trees; which of
them is flagged depends on the rest of the lint source list. With
scripts/lint.sh --ci and cppcheck 2.20.0 built from source:

- main itself fails on search_result_cmp ("file_order != 0 is always
  false");
- adding one more language grammar and its sources moves the finding
  to project_record_compare ("by_root != 0 is always false");
- carrying each comparator's order in one variable instead still fails,
  now as redundantAssignment on project_record_compare.

Both comparators now call nullable_strcmp(), which orders NULL like ""
exactly as the inline fallbacks did. lint-cppcheck passes again on
main and with the additional sources. Comparison order and results are
unchanged for every input.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Sarlette <282189207+SarletteD@users.noreply.github.com>
Vendors HeytalePazguato/tree-sitter-iec61131-3-st @ 00e24f50f8de (MIT,
(c) 2026 Jorge Centeno) unmodified as CBM_LANG_ST and registers it on
.st, with lang_specs node types for functions and methods, the
class-like POUs (FUNCTION_BLOCK, PROGRAM, INTERFACE, TYPE), properties,
calls, USING imports, branches, variables and assignments.

Vendoring: the grammar's LICENSE next to it, a THIRD_PARTY.md entry, a
MANIFEST.md paragraph and table row, regenerated
scripts/vendored-checksums.txt. Every language-count surface moves from
162 to 163 (README, docs/index.html, docs/llms.txt, pkg/npm/README.md,
the chocolatey nuspec, scripts/package-release.sh, server.json); four
places that still read 158 at the base (README.md, docs/index.html
twice, docs/llms.txt) are corrected as well.

Wired like the other languages: the language-name table in
userconfig.c, a grammar-regression fixture with its label golden, an
extension test, and the two ST call-node kinds plus the ST
routine-argument case in the call-node and language-registry repro
tests (the manifest assert message now states its real count).

Out of scope here: ST-specific definition and usage extraction, vendor
dialect normalization and the member layout of vendor exports; they
follow in separate changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Sarlette <282189207+SarletteD@users.noreply.github.com>
- extract_defs.c: st_base_classes() collects every base named by a
  repeated `extends`/`implements` field on a FUNCTION_BLOCK or
  INTERFACE (the generic first-match field probe keeps only one);
  class_base_names() dispatches ST there and every other language to
  the unchanged extract_base_classes().
  find_class_body() returns the declaration node itself for ST before
  the generic "body" field probe runs, so a trailing body statement
  tagged with that field name no longer hides the block's METHOD,
  PROPERTY and VAR members.
- extract_usages.c: occurrence_specs[CBM_LANG_ST] binds the names of a
  variable_declaration (its plural `names` field), so declared VAR
  names are bindings, not usages; a declaration's `initial_value` stays
  an ordinary read (is_non_binding_decl_field() now holds that check
  together with the existing `type` field check).
- tests: the ST label golden now counts the block's METHOD as a Method
  (it was a Function while the block body was misread); function-block
  members, members before a trailing body
  statement, interface members, a type alias keeping its name, declared
  names vs. type and initializer usages, and every EXTENDS/IMPLEMENTS
  base collected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Sarlette <282189207+SarletteD@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant