Reduce Windows CI runner time without dropping validation - #165
Merged
Merged
Conversation
Drop identical no-tool transcripts and duplicate model/argument witnesses, reuse a labeled comparison result for schema validation, and retain direct case-variant parsing and native failure-artifact checks. Keep the automatic Windows gate and every policy, isolation, timeout, ledger, skill, and unique parser boundary unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Run all Release test modules with ETW-sensitive CLI tests isolated; retain full generic capture checks on Linux and focus the Windows-native argv proof. Exercise pure fake-agent shapes through production validators and overlap its remaining isolated process contract with independent Windows fake-driven gates only after native ETW work completes. Local exact-source job phases: 601.10s baseline versus 282.86s candidate. Hosted Windows runtime and runner-price savings still require an authorized PR run. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The fake-agent deadline is not enforced during other gates, and cleanup contains a process-exit race.
Review effort: Balanced
Findings: 1
What changed in this PR
Reduces Windows CI duration while preserving Windows-specific ETW, test, capture, and evaluator coverage.
Changes:
- Parallelizes non-ETW test modules while isolating CLI/ETW tests.
- Overlaps fake-agent evaluation with focused Windows contracts.
- Replaces redundant fake-host processes with direct production-validator checks.
The hosted run passed in ~479 seconds, improving on 845 seconds but missing the ≤422-second target.
| File | Description |
|---|---|
tools/Test-WindowsDotNet.ps1 |
Partitions and validates Windows test modules. |
tools/Test-WindowsContracts.ps1 |
Orchestrates overlapping Windows gates. |
tools/Test-CaptureCommandTrace.ps1 |
Adds focused native-argv validation. |
tools/Test-AgentEval.ps1 |
Moves schema cases to direct validator tests. |
tools/fixtures/Fake-CopilotEvalHost.ps1 |
Removes redundant fake-host modes. |
eval/README.md |
Documents revised evaluation coverage. |
.github/workflows/ci.yml |
Integrates the optimized Windows workflow. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Monitor each owned foreground contract child in bounded intervals while checking the fake-agent deadline, including output drain. Stop only owned process trees on timeout or failure and tolerate natural exit racing with Kill. A one-second deadline now interrupts a running gate promptly; the normal seven-gate contract, forced step failure, and cleanup states remain verified locally. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
ciaggregate, job count, and 20-minute cap are unchanged.Evidence
No raw ETW captures or private investigation artifacts are included.
Hosted results and review follow-up
599e968: Windows 479s (7m59s), Linux ARM64 7m43s, agent-files 1m24s, and requiredcipassed; published-rate four-job cost $0.134.da73210: Windows 519s (8m39s), Linux ARM64 7m56s, agent-files 1m16s, and requiredcipassed; published-rate four-job cost $0.144. Relative to the 845s / $0.204 baseline, this head saves 326s (38.6%) and $0.060 per equivalent CI run. It misses the ≤422s half-time goal by 97s.Killrace.da73210monitors each owned foreground gate and output drain, interrupts an expired fake deadline, and makes cleanup idempotent. Locally a one-second deadline interrupted a running gate in 1.51s with both owned PIDs gone; forced foreground failure and the complete seven-gate path also passed.The second hosted run took 40s longer than the first (test step +10s, overlapped contracts +30s). That is an observed difference across hosted runs, not an isolated causal estimate of the review fix's overhead. No merge, release, or raw investigation trace is part of this PR.