Skip to content

Tracking: structural interpreter work to reach CRuby-interp parity (B1–B4) #384

Description

@linyiru

Part of #378. Each item gets its own issue, and B2 its own ADR, before implementation. The ordering will be re-ranked after the A items land and the per-request census is re-run.

  • B1: kwargs without a Hash.
    • Today every kwargs call does op_new_hash, then hash-literal dedup, then SipHash, then intern, and only then unpacks. That costs 370 ns versus 19 ns on CRuby.
    • CRuby passes keywords positionally via the call-info kw_arg list and allocates only for **opts.
  • B2: 8-byte Copy Value.
    • Value is a 16-byte enum holding Rc (Str, Class), so even an Int push/pop runs the clone/drop glue.
    • In the empty-loop profile, drop_glue, Value::clone and Vec::push account for about 25% and the Vm::run dispatch for about 20%. This is the root of the ~4× base step cost.
    • Candidates are NaN-boxing or a tagged pointer, with Str/Class moved behind heap ids. Needs an ADR; co-design with the C value representation.
  • B3: call-site cache holding the resolved method plus a call handler.
    • Skip the do_call cascade on a hit, along the lines of CRuby's vm_call_handler.
    • Replace &*name == "..." string compares and runtime intern on hot paths with precomputed SymIds.
  • B4: Proc-free yield and frame elision.
    • yield costs 224 ns versus 27 ns on CRuby.
    • ADR 0037 measured −21..34% per call for the frameless zero-arg tier; extend that approach.

Target: the Rails hello world bench within 1.5× of CRuby interp.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    researchNeeds analysis before implementation scoping

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions