Skip to content

Named arguments and ... notation - #57

Open
oflatt wants to merge 11 commits into
mainfrom
oflatt-named-args
Open

oflatt wants to merge 11 commits into
mainfrom
oflatt-named-args

Conversation

@oflatt

@oflatt oflatt commented Jul 14, 2026

Copy link
Copy Markdown
Member

No description provided.

@oflatt
oflatt marked this pull request as ready for review September 21, 2026 20:08
@oflatt
oflatt requested a review from a team as a code owner September 21, 2026 20:08
@oflatt
oflatt requested review from saulshanabrook and removed request for a team September 21, 2026 20:08
oflatt and others added 2 commits September 21, 2026 20:13
Resolve conflicts with main:

- `Cargo.toml`/`Cargo.lock`: take main's dependency set and pin egglog to
  upstream `9063586`, which already contains egglog#1008, so main's
  temporary `[patch]` onto the egg-smol fork is no longer needed. Drop the
  stale "local dev against the forked egglog" comment; the quasiquote work
  it referred to is not used by this branch.
- `src/lib.rs`: keep main's rewritten crate docs and module list, and add a
  named-arguments bullet under "Language and values".
- `CHANGELOG.md`: record named arguments under Unreleased/Added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`make nits` runs `cargo clippy --tests -- -D warnings` against a crate with
`#![warn(missing_docs)]`, so `NamedChange::delete` and `NamedChange::subsume`
need doc comments. Also apply `cargo fmt`, which the current rustfmt wants
for `named_args.rs` and for the module ordering the merge produced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

@saulshanabrook saulshanabrook left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TLDR: multiple ellipsis shadow names, doesnt work with include, and metadata not linked to push/pop

For push/pop, I believe I added some storage for extensions for named schedules to deal with this kind of thing and attach things correctly as an extension point you can look at

Review from codex:

Reviewed 207518e. I found three correctness issues worth fixing before merging.

  1. [P1] Ellipsis variables can collide and silently change query results — named_args.rs:145. The second fresh variable for field x has the same name as the first for x1. This program incorrectly fails; replacing each ... with _ _ passes:

    (relation R (:x i64 :x1 i64))
    (relation Hit ())
    (R 1 2)
    (rule ((R ...) (R ...)) ((Hit)))
    (run 1)
    (check (Hit))

    Use a fixed fresh-variable hint, as wildcard parsing does, rather than user-provided field names.

  2. [P2] Named calls fail after included declarations — named_args.rs:185. If defs.egg declares (relation R (:x i64 :y i64)), then (include "defs.egg") followed by (R :x 1 :y 2) fails with Unbound symbol :x. The caller is parsed before the include executes, so the call macro is registered too late. Included declarations need to be available when subsequent calls expand.

  3. [P2] Field metadata survives scoped redeclarations — named_args.rs:539. (push), a named one-argument declaration of R, (pop), then a positional two-argument declaration causes (R 1 2) to fail with “takes 1 argument.” Parsing retains the old macro, and positional declarations never replace it. Metadata needs to follow declaration scope and redeclaration.

For succinctness, I’d prioritize:

  • Drop the unrelated dependency upgrade and stale development comment: the parser APIs used here already exist at the previous revision.
  • Keep the macro implementation types private; expose register_named_args.
  • Share the duplicated named-field validation between schemas and datatype variants.
  • Inline register_named_call, which is a trivial forwarding helper.

Validation: All six added tests and 47 Rust integration tests pass. The file suite passes 44/45; the remaining math_backoff failure also reproduces on the base commit. Clippy passes. Formatting fails, and the committed lockfile references a different egglog revision than the manifest, preventing --locked builds.

The PR checkout is unchanged; no GitHub comments were posted.

All three correctness issues in the review share one cause: field names were
registered while parsing. egglog parses a whole program before running any of
it, so a parse-time macro sees declarations in source order, not execution
order.

Expanding in a `CommandMacro` fixes them together, because command macros run
after every preceding command has taken effect:

- A call now sees declarations pulled in by an earlier `(include ...)`, which
  had not been read yet when the caller was parsed.
- Field names are recorded per `(name, arity)`, and each call resolves against
  the arity the e-graph currently reports for that name, so a declaration that
  has been popped or redeclared can no longer rewrite calls.
- Each `...` binds fresh variables from a fixed hint instead of the field
  names, which could collide with each other ("x" at 1 and "x1" at 0 both
  gave "x1"). The pinned egglog also makes `SymbolGen` collision-free, so this
  no longer depends on that fix.

No parse-time macro is left: `:name` and `...` already parse as ordinary
variables, and `:name Sort` pairs already parse as a longer sort list, so the
declaration commands no longer need shadowing. That drops the `set`/`delete`/
`subsume` macros (their arguments are rewritten in place instead), makes every
implementation type private, and shares one field-name validator between
schemas and datatype variants. Named arguments are now registered on the
e-graph rather than the parser, since expansion needs the command stage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
oflatt and others added 2 commits September 22, 2026 11:20
Address review: expand named arguments during command macros
`main` gained `unstable-subst` (#60). The only conflict was two doc bullets
added at the same place in the crate docs; both are kept. Both sides already
pin the same egglog, so the manifest and lockfile merged cleanly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@oflatt-claude

Copy link
Copy Markdown
Contributor

Thanks for the detailed review — all three are fixed, in #72 (already merged into this branch) plus #73 for the fresh conflict with main after #60 landed.

They turned out to share one cause: field names were registered while parsing, and egglog parses a whole program before running any of it, so a parse-time macro only ever sees declarations in source order. Expansion now happens in a CommandMacro, which runs after each preceding command has taken effect, and that fixes all three together.

  1. Ellipsis collisions — each ... now mints fresh variables from a fixed hint, the way wildcard parsing does. Worth flagging: your repro already passes on the current pin without that change, since SymbolGen became collision-free upstream; the fixed hint just removes the dependency on that fix. Regression test in tests/named-args-ellipsis-fresh.egg.
  2. include — fixed by the same move; a call now sees declarations the include pulled in. tests/named-args-include.egg.
  3. push/pop — names are recorded per (name, arity), and every call resolves against the arity the e-graph currently reports for that name, so a popped or redeclared declaration can no longer rewrite calls. tests/named-args-scope.egg covers your repro plus redeclaration at the same arity.

I confirmed 2 and 3 fail on the old implementation and pass now.

On the extension storage you pointed at: extension_state is only reachable from a UserDefinedCommand, while CommandMacro::transform gets just symbol_gen and type_info. Since type_info is itself push/pop-scoped, resolving through it gives the same guarantee without new egglog API. The registry itself isn't scoped, but a stale entry is unreachable — it can only be found through an arity the live type info still reports. Happy to move it onto extension state if you'd rather have it explicit; it would need the macro to see the e-graph.

Succinctness:

  • No parse-time macro is left at all. :name and ... already parse as ordinary variables, and :name Sort pairs already parse as a longer sort list, so nothing needs shadowing. That also removed the set/delete/subsume macros (their arguments are rewritten in place) and register_named_call.
  • All implementation types are private; register_named_args is the only public item.
  • One split_fields validator is shared between schemas and datatype variants.
  • named_args.rs is 607 → 494 lines with the fixes included.
  • The dependency bump and the stale "local dev against the forked egglog" comment are gone — the branch just follows main's pin now, and the lockfile matches the manifest, so --locked builds work.

One API change to flag: register_named_args takes &mut EGraph instead of &mut Parser, since expansion needs the command stage. experimental_parser() on its own no longer provides named arguments.

Formatting is fixed too. make test and make nits both pass (172 tests, doctests, clippy with -D warnings, fmt --check).

@saulshanabrook saulshanabrook left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed d3255ef. The original three reproducers are fixed, but I found three regressions:

  1. [P1] Cloned e-graphs share mutable field mappings — named_args.rs:79. Clone an empty e-graph, then declare R(:a i64 :b i64) in one and R(:b i64 :a i64) in the other. In the first graph, (R :a 1 :b 2) now silently inserts (R 2 1). The command registry clones its Arc, sharing this mutable map. Metadata needs to follow the actual declaration and e-graph snapshot; (name, arity) is insufficient.

  2. [P2] Named calls inside extension commands bypass expansion — named_args.rs:214. The pinned Command::visit_exprs skips UserDefined commands, including experimental extract and run-schedule. This previously working program now fails with Unbound symbol :x:

    (datatype T (C :x i64))
    (extract (C :x 1))

    Explicitly traverse those commands’ expression arguments.

  3. [P2] Positional tables reject previously valid variables — named_args.rs:325. This succeeds on the previous head but fails now:

    (relation R (i64))
    (let :value 1)
    (R :value)

    Marker interpretation should remain limited to declarations with named fields; ordinary positional calls should pass through unchanged.

For succinctness, the rewrite removes substantial duplication. One remaining improvement: declare() could modify the command’s schema in place, combining constructor/function handling instead of reconstructing every unchanged field.

Validation: All 10 named-argument tests and 101 Rust integration tests pass. Formatting, Clippy, and locked builds now pass. The findings above have separate reproductions. No repository files or GitHub comments changed.

@oflatt-claude

Copy link
Copy Markdown
Contributor

Good catches — all three are fixed in #73, and the first one is the reason I moved the whole thing back into the parser.

[P1] cloned e-graphs. You're right that (name, arity) was never going to be enough. The deeper problem is that CommandMacro::transform only receives a SymbolGen and a TypeInfo, so there is no per-e-graph storage at that stage at all — the field names had to live in the macro object, whose Arc every clone shares. That's also why I couldn't use the extension state you pointed me at: extension_state is only reachable from a UserDefinedCommand.

The parser is the right home after all, for exactly the reason you gave: it's part of the e-graph, so it's cloned with one and snapshotted by push/pop, and a per-name expression macro only fires for the declaration that registered it. So expansion is back in the parser, and the two things the original version was missing are fixed head-on:

  • (include ...) is expanded while parsing instead of being left for egglog to read after the whole program is parsed, so an included declaration is registered before the calls that follow it.
  • A positional declaration registers a pass-through macro rather than nothing, shadowing field names registered earlier for that name. The thing that makes this sufficient is that egglog rejects redeclaring a bound name — (push) (relation R (i64)) (pop) over an existing R fails with Function already bound R — so the only way one name is declared twice is (push)/declare/(pop)/declare, and there the later declaration in the source is exactly the one that's live afterwards. Source order and declaration scope agree, which is what the parser needs.

[P2] extension commands and [P3] positional variables both disappear with the move rather than needing their own fixes: parse-time macros fire wherever parse_expr runs, including user-defined command arguments, and a macro only exists for names that were actually declared, so (R :value) on a positional relation is never inspected for markers. I dropped the speculative "not declared with named fields" error that caused the latter — that was me guessing, and you were right that it shouldn't.

There are now regression tests for all six reported cases, and I checked each one fails on the implementation it was reported against:

Case Test
Ellipsis collisions tests/named-args-ellipsis-fresh.egg
Included declarations tests/named-args-include.egg
Scoped redeclaration tests/named-args-scope.egg
Cloned e-graphs test_named_args_do_not_leak_between_cloned_egraphs
Extension commands tests/named-args-extract.egg
Positional variables tests/named-args-positional-vars.egg

The succinctness changes from last round are kept: register_named_args is the only public item, one read_named_field validator is shared between schemas and variants, and register_named_call is gone. Your declare() note no longer applies, since that function belonged to the command-macro version.

One thing worth your judgement: expanding include at parse time means Command::Include no longer reaches egglog, and the file is read while parsing rather than while running. Same commands in the same order, but it is a visible change to how includes behave for anyone using experimental_parser(). If you'd rather not have that, the alternative is an egglog affordance — e.g. passing &mut extension_state alongside symbol_gen and type_info in CommandMacroRegistry::apply, which the borrow checker allows since they're disjoint fields — and I'm happy to do that instead.

make test and make nits pass: 175 tests, doctests, clippy with -D warnings, fmt --check.

@oflatt-claude

Copy link
Copy Markdown
Contributor

Correction to my last comment: include no longer needs the behaviour change I flagged, so disregard that question.

Only the registration has to move earlier, not the command. include is now read once while parsing, purely so the declarations inside register their field names before the calls that follow it are parsed — then those commands are discarded and Command::Include is passed through untouched, so egglog reads and runs the file exactly as it always has. The read is best-effort and never reports: a missing or unparsable file is left for egglog to fail on from its own include handling, so this can only ever add registrations. A missing include still fails with egglog's own IO error: ... No such file or directory.

The cost is that an included file is read twice, which seems a fair price for leaving include semantics alone. tests/named-args-include.egg covers both halves — the included relation is usable by name, and it is declared exactly once, since a second execution of the include would fail with Function already bound.

Pushed to #73. Still 175 tests, doctests, clippy and fmt clean.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants