Repository navigation
Conversation
Replace InferenceOptions.includeLogits: Bool with two orthogonal axes —
tokens: TokenRequest {.none, .sample} and logits: LogitsRequest
{.none, .lastPosition, .allPositions} — plus presets (.prefill / .extend /
.eval / .guided / .sampleWithLogits). Default (.sample, .none) equals the old
includeLogits: false, so decode is byte-identical. All call sites migrated to
presets; engines derive returnsLogits = (logits != .none); the pipelined engine
rejects logits != .none (was includeLogits).
132916f to
9431478
Compare
…binations Add an InferenceOptions.returnsLogits computed property (equivalent to logits != .none) and route all engines (sequential, static-shape, VLM, pipelined guard) and the MockEngine through it instead of open-coding the comparison. Document on the raw initializer that the (tokens, logits) space is intentionally permissive with no invalid combinations, that the presets cover the meaningful pairings, and that (tokens: .sample, logits: .allPositions) is the one representable pair without a dedicated preset. Pure readability refactor; behavior unchanged.
| /// `eval`, `guided`, and `sampleWithLogits` presets below cover all the | ||
| /// meaningful pairings; prefer them to constructing options by hand. The | ||
| /// only representable pair without a dedicated preset is | ||
| /// `(tokens: .sample, logits: .allPositions)` — sampling a token while also |
There was a problem hiding this comment.
grepped for .tokens, .logits, .allPositions and .lastPosition in swift/ and internal/. Outside InferenceEngine.swift, the only readers are the tests. Every engine (sequential, VLM, static-shape, pipelined) reads returnsLogits, which is just logits != .none.
So tokens: .none and .allPositions change nothing at runtime, so these are declared but no engine reads them 🤔
There was a problem hiding this comment.
Yes, this will be handled in a follow up.
(The engines do extra work today)
| /// Max tokens to generate. Nil = until EOS or context limit. | ||
| public var maxTokens: Int? | ||
| /// Include raw logits in each `InferenceOutput`. May incur GPU→CPU copy cost. | ||
| public var includeLogits: Bool |
There was a problem hiding this comment.
Public API breaking when we remove this
There was a problem hiding this comment.
let's mark this deprecated
|
|
||
| /// Presets for the common (tokens, logits) pairs. | ||
| extension InferenceOptions { | ||
| /// Warm the KV cache; no token or logits generation. |
There was a problem hiding this comment.
.prefill is documented as "no token or logits generation" but an engine still samples even though .guided and .eval set tokens: .none
There was a problem hiding this comment.
Updating documentation
Reintroduce the old InferenceOptions.includeLogits boolean as a deprecated initializer parameter and computed property that bridge to the (tokens, logits) axes: true -> logits: .lastPosition, false -> .none, with tokens still sampled. This avoids a hard public-API break from splitting the bool into the orthogonal axes — external callers keep compiling, with a deprecation warning pointing to the new API. Addresses review feedback.
Reword the .prefill preset from "no token or logits generation" to describe what the request needs. The (tokens, logits) axes express intent, and engines do not yet gate sampling on tokens: .none, so the old wording over-claimed runtime behavior. Addresses review feedback.
Summary
A code-health cleanup of the inference layer. It replaces the
InferenceOptions.includeLogits: Boolflag with two orthogonal request axes, so all four engines — sequential, pipelined, static/ANE, and VLM — describe a forward step the same way:tokens: TokenRequest(.none/.sample) — should the engine sample and return a token?logits: LogitsRequest(.none/.lastPosition/.allPositions) — which positions' logits should it surface?A few presets cover the everyday combinations:
tokenslogits.prefill.none.none.extend(alias.default).sample.none.eval.none.allPositions.guided.none.lastPosition.sampleWithLogits.sample.lastPositionMotivation
The single boolean had to answer two unrelated questions — whether to sample a token, and which logits (if any) to return — and it couldn't answer either well:
Splitting the request into an action axis and an output axis lets a caller ask for exactly what it needs, and lets an engine turn down a request it can't serve up front rather than fail partway through a forward pass.
Compatibility
The default,
(tokens: .sample, logits: .none), is equivalent to the oldincludeLogits: false, so decode is unchanged. Every call site now uses a preset, and engines derivereturnsLogits = (logits != .none).It does remove the public
includeLogitsproperty, which is a breaking change; the replacements are the two enums, thetokensandlogitsproperties, and the preset factories.Verification
This is a behavior-preserving refactor, and the output is byte-identical to the base commit:
Follow-ups
Intentionally left for separate changes:
.lastPositionand.none— a decode-throughput win the sequential and VLM paths don't take yet.tokensis.noneon the decode path..allPositionson the static and pipelined engines.