A small, agent-shaped CLI and MCP server for analyzing .NET CPU, allocation,
blocking, and wall-clock traces. Built on the
Microsoft.Diagnostics.Tracing.TraceEvent library; reads EventPipe
(.nettrace / .speedscope.json) and ETW (.etl) captures from both .NET and
.NET Framework runs.
filtrace targets .NET 10. Both heads are published on NuGet.org:
KlutzyNinja.Filtrace
(the filtrace CLI) and
KlutzyNinja.Filtrace.Mcp
(the MCP server).
Installing the CLI as a .NET global tool needs the .NET 10 SDK (dotnet tool
ships with it):
dotnet tool install --global KlutzyNinja.Filtrace
filtrace rank app.nettrace --metric cpuUpdate or remove it later with dotnet tool update --global KlutzyNinja.Filtrace
or dotnet tool uninstall --global KlutzyNinja.Filtrace.
The MCP server runs on demand - there is no install step. Add the stdio server to
your agent's MCP config and dnx (bundled with the .NET 10 SDK) fetches
KlutzyNinja.Filtrace.Mcp and launches it. See
Using filtrace from an AI agent for the exact
config block and the tool workflow.
Activate the current checkout for one Git repository without changing a global Filtrace installation:
./tools/Use-LocalFiltrace.ps1 -Action Install -TargetRepository ../consumer -Configuration ReleaseRun Install again to refresh the private CLI, MCP server, and skill from the
current source. -TargetRepository defaults to the current directory. Restore
the target's recorded MCP and skill baseline and remove the private CLI with:
./tools/Use-LocalFiltrace.ps1 -Action Restore -TargetRepository ../consumer -Configuration ReleaseThe command requires PowerShell 5.1 or 7, Git, and the .NET 10 SDK. It does not need elevation, change a global tool installation, or upload the prepared package.
Every analysis command takes a trace path and prints a dense text report (or compact
JSON with --format json); collect instead launches the executable it records.
The canonical investigation is orient -> rank -> drill ->
compare: inspect the capture, rank the matching metric, drill an unwindowed CPU
ranking when needed, then diff comparable traces or capture manifests against a baseline.
# Workflow: orient, rank the hottest frames, drill into one, then diff two runs.
filtrace info app.nettrace # 0. orient: format, providers, event counts, symbol rate
filtrace rank app.nettrace --metric cpu # 1. what's hot (self weight: milliseconds or raw samples)
filtrace callers app.nettrace MyApp.Parse # 2. who calls the hot frame
filtrace source app.nettrace --view lines --symbols bin/Release/net10.0 # 3. hot source lines
filtrace diff before.nettrace after.nettrace # 4. what changed between runs
filtrace batch BenchmarkDotNet.Artifacts/filtrace-runs/run/manifest.jsonThe same analysis core is exposed as a stdio MCP server: every analysis command has a
matching trace_* tool (eighteen in all - info -> trace_info, rank ->
trace_rank, callers -> trace_callers, and so on), returning the same envelope
shape and results. The capture and housekeeping commands (collect, cache) are
CLI-only, as is widening an ETW analysis with --all-processes; MCP
supports automatic or named-process scope. See
Using filtrace from an AI agent for the client
config and tool workflow.
Orient - see what a capture holds before ranking (the CLI counterpart of the
trace_info tool):
| Command | Purpose | Example |
|---|---|---|
info |
Format, sample count, lost-event and symbol-quality warnings, supported analyses, and per-analysis capture/event state | filtrace info app.nettrace |
Ranking - rank stacks by a metric:
| Command | What it ranks | Example |
|---|---|---|
rank |
Any metric (cpu, alloc, exceptions, threadtime, contention, wait, activity) |
filtrace rank app.nettrace --metric contention |
Implemented scope inventory:
- Named process: CLI
info,rank,source,callers,tree,classify,timeline,diff,batch,export, andreport(for--kind diskio); MCPtrace_info,trace_rank,trace_callers,trace_lines,trace_heatmap,trace_tree,trace_classify,trace_timeline,trace_diff,trace_batch, andtrace_export, plustrace_diskio. Every listed surface except the disk report auto-scopes a multi-process.etlto the busiest process tree. Runprocesses/trace_processesfirst to inspect the capture, then set--process <name>/processto override. CLI commands expose--all-processeswhere an aggregate is supported; MCPtrace_infoandtrace_rankexposeallProcesses. Capture-wideinfo/trace_infopreserves whole-capture totals but budget-limits its busiest-thread detail with an explicit truncation warning. Process inventory aggregates ETL and EventPipe CPU ownership without materializing frames; speedscope uses its full stack reader. Useinfo/trace_infofor frame-name and source/PDB quality. Other stack-backed MCP analyses have no all-process aggregate;trace_diskiois the exception and, like the CLI disk report, remains machine-wide by default. When scoped, disk reports correlateDiskIOInitissuer IRPs to completions because a completion PID may be System or Idle. This is direct issuer scope, not causal ownership: deferred file-system/cache write-back issued by System is included only when System is selected, while Idle-issued I/O appears only in an unscoped report. - Exact process ids: the same commands and tools accept
--pid <id>[,<id>](comma-separated, not repeated) /pidinstead of a name. A name substring is right for discovery, but a common host name such asdotnetmatches every unrelated instance in a machine-wide capture and ranks them together; an exact id set cannot. Prefer it for manifests and automation. The three selectors are mutually exclusive, an id reused by two processes in one trace is refused rather than merged, and an id that is not in the trace is reported. Fordiffbetween two.etltraces with different process ids, supply both--before-pid <id>[,<id>]and--after-pid <id>[,<id>](MCP:beforePidandafterPid). These scope the baseline and current traces separately, cannot be combined with the shared--pid,--process, or--all-processesselectors, and fail if either trace lacks a requested id. Manifest diffs instead use each case's recorded invocation ids; per-arm ids are not accepted for manifests. - Descendants: the same commands and tools accept
--children include|exclude/children. Both selectors follow descendants by default, because the common capture shapes put the measured work in a child the host launched. Passexcludeto separate a parent's own CPU from a child runtime's; without it a native host's own cost is blended with the CoreCLR frames of the child it launched. - Disk completion time: CLI
report --kind diskioand MCPtrace_diskioaccept--time <start>,<end>/time. The window filters completion timestamps after issuer-process IRP correlation; either bound may be open. - Invocation roots: CLI
lifecycleand MCPtrace_lifecycletake the same--process/--pidselectors, but each matched process instance is one invocation and descendants always follow, so neither takes--childrenor--all-processes. - Root subtree: CLI
rank,callers,tree,classify,diff,batch, andexport; MCPtrace_rank,trace_callers,trace_tree,trace_classify,trace_diff,trace_batch, andtrace_export. Set--root <frame>/rootto keep the subtree under a frame. Root filtering is stack ancestry, not causal correlation: stacks without the selected frame are excluded, including sibling workers. Root-aware structured results identifyrootKind: stackAncestryand report available versus retained weight and record counts; direct diffs report both sides, and manifest batch/diff report each case. Use an instrumented activity or validated time window for a parallel phase, and ETWthreadtimewhen sampled CPU does not explain elapsed time. - BenchmarkDotNet workload: CLI
rank,callers,tree,classify,diff,batch, andexportaccept--benchmark; MCPtrace_rank,trace_callers,trace_tree,trace_classify,trace_diff,trace_batch, andtrace_exportacceptbenchmark: true. The preset isolates theWorkloadActionsubtree from harness and overhead scaffolding; it is mutually exclusive with an explicit root.sourceviews are not root-aware, so narrow them by method/file and treat percentages as process-scoped whole-trace values.
The rank command adds two more scopes: --activity <name> (the CPU samples taken inside one
start-stop request/job) and --time <start>,<end> (milliseconds from the trace
start, either bound optional; any metric on .nettrace / .etl), to zoom in on one
request or the slice around a latency spike. Speedscope input is aggregate-only for
--time and warns that the window was ignored.
filtrace rank bdn.nettrace --metric cpu --benchmark # just the [Benchmark] code
filtrace rank bdn.nettrace --metric alloc --benchmark # allocations under the workload
filtrace processes machinewide.etl # list every process by weight
filtrace rank machinewide.etl --metric cpu --process MyApp # one process tree
filtrace rank machinewide.etl --metric cpu --pid 9144,40356 --children exclude # exactly those, parent-only
filtrace rank app.nettrace --time 1000,5000 # just the spike windowNative runtime symbols. Managed frames (including NGEN and ReadyToRun
framework methods) resolve for free from the trace's CLR rundown. The unmanaged
runtime frames - the GC, the JIT, memset / memcpy, write barriers - need PDBs
from the Microsoft public symbol server, which rank fetches only when you
opt in with --native-symbols (cached under --symbol-cache, default in the temp
path). It is off by default so analysis stays offline and deterministic; the first
run downloads, later runs hit the cache.
filtrace rank app.etl --metric cpu --process MyApp --native-symbols # name the GC/JIT/memcpy framesCPU drill-down - follow an unwindowed CPU ranking into detail:
| Command | Purpose | Example |
|---|---|---|
callers |
Immediate CPU callers of a frame, or a caller/callee view with --callees |
filtrace callers app.nettrace MyApp.Parse --callees |
source |
Ranked source lines or per-file heat | filtrace source app.nettrace --view heatmap --file Parser.cs |
tree |
Top-down CPU call tree from the root | filtrace tree app.nettrace --max-depth 5 |
Inventory - see what a (possibly machine-wide) capture contains:
| Command | Purpose | Example |
|---|---|---|
processes |
List processes by CPU-sample weight, to pick a --process target |
filtrace processes machinewide.etl |
classify |
Summarize CPU weight by runtime work category (ms when established, otherwise samples) |
filtrace classify app.etl --native-symbols |
Temporal - see what happened when, then scope a ranking to the busy window:
| Command | Purpose | Example |
|---|---|---|
timeline |
Aligned activity buckets or one bounded cross-lane snapshot | filtrace timeline app.nettrace --mode snapshot --at 1500 |
Compare and export:
| Command | Purpose | Example |
|---|---|---|
diff |
Absolute/normalized CPU changes for traces or paired manifests | filtrace diff before.nettrace after.nettrace |
batch |
One compact ranking across every capture-manifest case | filtrace batch run/manifest.json |
export |
Write a flame graph (speedscope / chromium) | filtrace export app.nettrace --format speedscope -o app.json |
Structured reports:
| Command | Purpose | Example |
|---|---|---|
report |
GC, JIT, thread-pool, or physical disk-I/O report | filtrace report app.nettrace --kind gc |
lifecycle |
Per-invocation wall-clock phases: root lifetime, first child, child span, teardown (ETW) | filtrace lifecycle run.etl --process myapp --image hostfxr |
events |
Query raw events, filtered by name / payload / pid / tid, paged | filtrace events app.etl --payload ConnectionReset |
Capture (Windows, elevated) - record an ETW .etl yourself, no external recorder:
| Command | Purpose | Example |
|---|---|---|
collect |
Launch an executable and record a CPU, thread-time, startup, or physical disk-I/O .etl |
filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl --profile threadtime |
filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl # CPU
filtrace collect --launch dotnet --launch-args MyApp.dll --output tt.etl --profile threadtime
filtrace collect --launch MyApp.exe --output start.etl --profile startup # low perturbation
filtrace collect --launch MyApp.exe --output io.etl --profile diskio # physical disk/files
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --working-directory C:\src\app --output build.etl
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --output build.etl --rundown --rundown-pid 1234 # known persistent CLR server
filtrace collect --launch MyApp.exe --output ring.etl --max-size-mb 512 # bounded ring buffer--working-directory sets and records the absolute directory inherited by every
subject launch. Omit it to inherit the collector's current directory.
--rundown appends and merges a separate minimal CLR naming rundown after the
launched command exits. Use it only when captured CPU belongs to managed servers
that were already running and remain alive, such as compiler/build servers. The
opt-in pass is machine-wide by default; pass up to 8 comma-separated exact ids
to --rundown-pid to filter CLR naming events when the target servers are already
known. The capture result records that filter. Rundown requests 512 MB of ETW
buffers, and TraceEvent's derived maximum-buffer count permits the pool to grow
to roughly 641 MiB. Provider activation can take up to 30 seconds, followed by up
to 30 seconds of quiescence polling; the subsequent ETL merge has no timeout.
Rundown can add hundreds of megabytes. It is not valid with --profile diskio or
--max-size-mb; inspect the capture and info lost-event warnings before
trusting resolved names.
A targeted process must be running before capture and remain the same process
through rundown; exit or PID reuse fails explicitly.
The diskio profile enables only physical DiskIO/DiskIOInit events, process/thread
attribution, and the DiskFileIO name rundown. It deliberately omits CPU sampling,
stacks, verbose FileIO, and CLR events. The capture and unscoped disk report are
machine-wide; write the trace to a different volume when recorder writes must not
contend with the workload volume. Scope a report to direct issuers with
--process or --pid, and to completion time with --time.
With --format json, stdout contains only the capture-result JSON; identified
subject stdout and stderr are forwarded to stderr. The command still exits successfully
when capture succeeds even if the subject fails, and reports that failure in
processExitCode and invocations. Text output continues to inherit the subject's streams.
For an EventPipe (.nettrace) capture - cross-platform, no elevation - use the
first-party dotnet-trace (dotnet tool install -g dotnet-trace, then
dotnet-trace collect -- <app>); collect is ETW-only.
File ops - manage the ETLX conversion cache filtrace keeps beside a trace:
| Command | Purpose | Example |
|---|---|---|
cache |
Build/reuse or remove the ETLX cache | filtrace cache app.nettrace --action convert |
ETLX conversion is coordinated per canonical trace path across threads and
processes, with unique temporary files and atomic publication. Filtrace records
the converter epoch, backend binary and source/cache file facts in
<ETLX path>.filtrace.json (distinct from capture metadata at
<trace path>.filtrace.json); a cache hit checks these small facts, not a hash
of the entire trace. A readable ETLX with no valid marker is not reused or
overwritten: run filtrace cache app.nettrace --action clean to explicitly
discard it, then convert or analyze again. This one-time rebuild can take
seconds for large traces. Owned caches rebuild when the source changes, and
interrupted publications recover. clean removes both ETLX and its marker.
Same-trace MCP queries may run in parallel; trace_info.etlxCacheState and
cache --action convert report hit, waited, converted, or recovered.
The previous command names remain callable during the current migration window and print their canonical replacement to stderr, but they are hidden from top-level help and are not used in examples or generated guidance. Removal requires the explicit VN5 migration policy; it is not tied automatically to the passage of one preview release:
| Previous names | Canonical command |
|---|---|
cpu, alloc, exceptions, threadtime |
rank --metric <name> |
lines, heatmap |
source --view <name> |
gcstats, jitstats, threadpool, diskio |
`report --kind gc |
convert, clean |
`cache --action convert |
Run filtrace <command> --help for the full option set of any command.
filtrace is built for an agent mid-investigation. Two ways to wire it in:
-
MCP server - add the stdio server so the agent calls the
trace_*tools directly:{ "servers": { "filtrace": { "type": "stdio", "command": "dnx", "args": ["KlutzyNinja.Filtrace.Mcp", "--yes"] } } } -
CLI - install the global tool (
dotnet tool install -g KlutzyNinja.Filtrace) and let the agent shell out tofiltrace <verb>.
Either way, the canonical loop is orient -> rank -> drill -> compare: read
trace_info (CLI: filtrace info) first; when symbol resolution is below 0.8,
inspect its warning and unresolved rows. Treat that as frame-name quality; before
source-line analysis, inspect sourceResolution for exact matching PDB modules,
mapped sampled managed frames, searched directories, highest-unmapped modules, and
pdbIdentityMismatchModules. A mismatch means a same-named local PDB was found but
its GUID or age differs from the trace. Once the relevant module matches, use
sourceMappedManagedMethodCount versus sampledManagedMethodCount to confirm sampled
methods resolve sequence points; unmappedNamedManagedFrameCount and
highestUnmappedMethods expose the remaining <no source> impact.
Use the generated BenchmarkDotNet child output when the outer build PDB does not
match; use native symbols for CPU ETW runtime frames as applicable. Rank by the metric that matches the question (cpu, alloc, exceptions,
threadtime, contention, wait, activity); for an unwindowed CPU ranking, drill the
hot frame with callers / lines / tree; diff comparable CPU traces against a baseline.
| Path | Purpose |
|---|---|
src/Filtrace.Core/ |
Analysis core: trace readers, stack-source providers, the provider-agnostic question-service engine. The only place logic lives. |
src/Filtrace/ |
CLI host, packaged as the filtrace .NET global tool. |
src/Filtrace.Mcp/ |
Stdio MCP host, packaged separately for dnx KlutzyNinja.Filtrace.Mcp. |
benchmarks/Filtrace.Benchmarks/ |
BenchmarkDotNet performance harness for the analysis core. |
benchmarks/Filtrace.PerfWorkload/ |
Parameterized CPU/activity workload for reproducible Track D traces. |
tests/Filtrace.Core.Tests/ |
Unit + golden-file contract tests. |
tests/Filtrace.Parity.Tests/ |
Numeric parity against the frozen legacy oracles. |
eval/ |
Headless-agent eval harness, tasks, baselines. |
docs/ |
Design, roadmap, competitive analysis, and the single-source workflow text for the skill / README / help. |
.agents/skills/filtrace/ |
The shipped agent skill. |
filtrace carries its own Directory.Build.props, Directory.Build.targets,
Directory.Packages.props, global.json, and .editorconfig (root = true),
so the build is fully self-contained. Its only external dependency is the
published KlutzyNinja.Touki NuGet package; it references no other project.
cd filtrace
dotnet build filtrace.slnx
dotnet test filtrace.slnx