10 KiB
Build Plan — Jai Language Plugin for IntelliJ
Design constraint: every milestone must be verifiable by an agent from the command line. No manual IDE launch, no human eyeballing. This drives the architecture below, not just the CI config.
Companion doc: docs/JAI_LANGUAGE_REFERENCE.md (language facts, derived from
~/.local/jai).
0. Verified starting conditions
Checked empirically before writing this plan:
| Fact | Status |
|---|---|
| Scaffold is the bare JetBrains template (demo tool window) | confirmed |
| Gradle 9.6.1, IntelliJ Platform 2025.3.5, Kotlin 2.3.20, IPGP 2.18.1 | confirmed |
./gradlew help succeeds |
confirmed |
No JDK on PATH; JAVA_HOME unset |
⚠️ blocker |
Adoptium JDK 21.0.12 present at ~/.gradle/jdks/… |
confirmed working |
org.gradle.java.home in gradle.properties does not fix it |
confirmed — gradlew needs a JVM to launch itself |
| Wrapper script that auto-discovers the JDK does fix it | confirmed, BUILD SUCCESSFUL |
generateLexer / generateParser available via bundled org.jetbrains.intellij.platform.grammarkit |
confirmed by gradlew tasks --all |
| Not a git repository | ⚠️ blocker for safe agent iteration |
Jai corpus: 714 .jai files, 16.4 MB |
confirmed |
The 714-file corpus is the single most important asset for this project — see §3.
1. Phase 0 — Make the project agent-operable (do this first)
Nothing else is testable until this is done.
./jaigradle— committed wrapper that resolves a JDK then delegates to./gradlew. Verified working. Every later instruction uses this, so an agent never has to think aboutJAVA_HOME.git init+.gitignoreforbuild/,.gradle/,.intellijPlatform/, and generated sources. Agents need a clean diff and a rollback point.AGENTS.md— the contract for future agents: the one command to run (./jaigradle check), where the corpus lives, where generated code lives and that it must never be hand-edited, and a pointer to the language reference.- Strip the template — delete
MyToolWindowFactory, the demo tool window registration, and the message bundle stub; rewriteplugin.xmlmetadata. - Prove the harness — add one trivial passing test and confirm
./jaigradle testruns it headlessly. This validates the whole loop before any real work.
Exit gate: ./jaigradle check green, from a shell with no JAVA_HOME.
2. Architecture decisions (chosen for testability)
2.1 Hand-written lexer, not JFlex — recommended
Standard practice is JFlex, but two Jai features argue against it:
- Here-strings have a dynamic terminator.
#string DONE … DONEwhere the terminator is an arbitrary identifier captured at runtime. JFlex is a DFA generator; it cannot match a token captured earlier. This is only expressible as hand-written Java inside a JFlex action — i.e. the hard part isn't in the grammar anyway. - Nested block comments need a depth counter, again custom action code.
Add the practical argument: the lexer is the component agents will iterate on
most, and a hand-written LexerBase removes a codegen round-trip from that
loop. It is also directly unit-testable with plain JUnit.
I'd write JaiLexer extends LexerBase by hand, with JaiTokenTypes as the
shared token holder. If this turns out badly we can retreat to JFlex later —
the token set and tests stay valid either way, so the decision is reversible.
2.2 Grammar-Kit BNF for parser + PSI — recommended
Here codegen genuinely pays: Grammar-Kit generates the PSI class hierarchy,
which is a large amount of boilerplate. It's the standard path, and I verified
the tasks run headlessly from our existing plugin (no extra dependency — the ID
is org.jetbrains.intellij.platform.grammarkit, bundled since IPGP 2.12.0).
The .bnf references the hand-written lexer's token types. Generated sources go
to a dedicated src/main/gen root, git-tracked so diffs are reviewable, and
regenerated by ./jaigradle generateParser.
Caveat: Grammar-Kit does not support two-pass generation, so no method mixins —
use the mixin/extends attributes instead.
2.3 Sequencing: ship a useful plugin before touching the parser
Syntax highlighting needs only the lexer. That's most of the perceived value and it de-risks the schedule. The parser is Phase 3, not Phase 1.
3. The testing strategy
This is the part that answers "agents test it without me booting an IDE."
Tier 0 — Corpus invariants (highest value, no IDE, plain JUnit)
Run the lexer across all 714 real .jai files and assert:
- Round-trip: concatenating every token's text reproduces the source file byte for byte. An IntelliJ lexer must tile the input with no gaps or overlaps; violating this corrupts the editor. Over 16 MB of real code this one assertion catches the overwhelming majority of lexer bugs.
- No
BAD_CHARACTERtokens anywhere in the distribution. - Progress: every
advance()strictly increases the offset — catches infinite loops, which otherwise hang the IDE rather than failing visibly.
This is the flagship gate. It's fast, deterministic, needs no IntelliJ fixture, and it is real-world coverage rather than toy snippets. Failures should report file, line, and column.
Tier 1 — Golden token dumps (LexerTestCase)
Hand-picked adversarial snippets, one per gotcha in the reference doc §14:
nested comments, here-strings, --- vs -- vs -> vs -=, ,,, .{/.[/
.IDENT, #char "a" (and that ' is not a delimiter), == before {,
hex/binary/hexfloat/underscored numerics, backticked identifiers.
Golden files make regressions obvious in a diff.
Tier 2 — Parsing tests (ParsingTestCase)
Fixture .jai files with expected PSI trees in .txt. Useful property for
agents: ParsingTestCase writes the expected file automatically if missing,
so adding a case is cheap — but the generated tree must be reviewed before
committing, or the test asserts nothing.
Tier 3 — Corpus parse gate
Parse all 714 files, assert zero PsiErrorElement. This is the acceptance
criterion for "the grammar is correct," and no human inspection can match it.
Expect to allow a small, explicit, shrinking allowlist of known-unsupported files rather than blocking the phase on 100% — but the allowlist must be committed and visible so it can't quietly grow.
Tier 4 — Code-insight tests (BasePlatformTestCase, headless)
myFixture drives annotators, folding, brace matching, commenter, completion,
formatter, and rename entirely headlessly. Highlighting is verified with
checkHighlighting against <warning>-annotated fixtures — this is how you
check colors without looking at a screen.
Tier 5 — Packaging and compatibility
./jaigradle verifyPlugin (Plugin Verifier — catches API misuse and
compatibility breaks) and buildPlugin (produces the installable ZIP).
Single entry point
./jaigradle check runs Tiers 0–4. Agents run one command; runIde exists for
you but is never required of an agent.
4. Phased milestones
Each phase has a machine-checkable gate. Do not advance without a green gate.
| Phase | Deliverable | Gate |
|---|---|---|
| 0 | Env fix, git, AGENTS.md, template stripped | ./jaigradle check green with no JAVA_HOME |
| 1 | File type, icon, JaiTokenTypes, hand-written lexer |
Tier 0 corpus invariants pass on all 714 files |
| 2 | SyntaxHighlighter, color settings page, commenter, brace matcher |
Tier 1 golden dumps; highlighter maps every token type |
| 3 | .bnf grammar, generated parser + PSI, ParserDefinition |
Tier 2 golden trees; Tier 3 corpus parse ≥ target — done, 100.0% (714/714) |
| 4 | Structure view, folding, #import/#load reference resolution + go-to-definition |
Tier 4 fixture tests — done |
| 5 | Completion (keywords, directives, module names), rename, find-usages | Tier 4 fixture tests — done (70-test suite) |
| 6 | Formatter, code style settings | Formatter round-trip: formatting the corpus is idempotent |
| 7 | Inspections (e.g. #must misuse), quick fixes, live templates |
Tier 4 + verifyPlugin |
| 8 | Optional: run-configuration to invoke the jai compiler, parse its error output |
Integration test against ~/.local/jai/bin |
Phase 6's idempotence check (format twice, assert no change) is another corpus-scale invariant that needs no human judgment.
5. Risks and how the plan handles them
| Risk | Mitigation |
|---|---|
| Environment is the real blocker — no JDK on PATH | Phase 0 wrapper script, already verified |
| Grammar-Kit codegen needing the IDE UI | Verified false — tasks run from CLI |
| Here-strings breaking a DFA lexer | Drove the hand-written lexer decision |
Nested comments / ' / .{ mis-lexed |
Explicitly enumerated as Tier 1 cases |
| Grammar churn silently breaking the parser | Tier 3 corpus gate |
| Generated sources hand-edited then lost | Committed to src/main/gen, AGENTS.md forbids editing, regeneration is one command |
| Slow feedback discouraging test runs | Tier 0 is plain JUnit with no fixture boot; keep it under a few seconds |
| Corpus files using features we never support | Explicit committed allowlist, not a silently loosened assertion |
6. Open questions for you
- Scope/ambition — stop at the current syntax, navigation, and basic refactoring support (Phases 0–5), or go all the way to formatter and inspections (0–7)?
- Compiler integration (Phase 8) — worth it? It's the only phase needing
the actual
jaibinary, and it's the least testable. - Target IDE — IntelliJ IDEA only (current setting), or all JetBrains IDEs?
The latter is mostly a
plugin.xmldependency change, cheapest decided now. - JDK policy — rely on the Gradle-provisioned JDK via the wrapper (zero setup, works today), or install a system JDK via Homebrew? I'd default to the wrapper.
Phases 0–5 are now implemented and verified headlessly. The remaining decision is whether to continue with the optional formatter, inspections, and compiler integration work in Phases 6–8.