Skip to content

Repository files navigation

filtrace

CI KlutzyNinja.Filtrace KlutzyNinja.Filtrace.Mcp License: MIT

A small, agent-shaped CLI and MCP server for analyzing .NET CPU, allocation, blocking, and wall-clock traces. Built on the Microsoft.Diagnostics.Tracing.TraceEvent library; reads EventPipe (.nettrace / .speedscope.json) and ETW (.etl) captures from both .NET and .NET Framework runs.

Install

filtrace targets .NET 10. Both heads are published on NuGet.org: KlutzyNinja.Filtrace (the filtrace CLI) and KlutzyNinja.Filtrace.Mcp (the MCP server).

CLI (global tool)

Installing the CLI as a .NET global tool needs the .NET 10 SDK (dotnet tool ships with it):

dotnet tool install --global KlutzyNinja.Filtrace
filtrace rank app.nettrace --metric cpu

Update or remove it later with dotnet tool update --global KlutzyNinja.Filtrace or dotnet tool uninstall --global KlutzyNinja.Filtrace.

MCP server

The MCP server runs on demand - there is no install step. Add the stdio server to your agent's MCP config and dnx (bundled with the .NET 10 SDK) fetches KlutzyNinja.Filtrace.Mcp and launches it. See Using filtrace from an AI agent for the exact config block and the tool workflow.

From source

Activate the current checkout for one Git repository without changing a global Filtrace installation:

./tools/Use-LocalFiltrace.ps1 -Action Install -TargetRepository ../consumer -Configuration Release

Run Install again to refresh the private CLI, MCP server, and skill from the current source. -TargetRepository defaults to the current directory. Restore the target's recorded MCP and skill baseline and remove the private CLI with:

./tools/Use-LocalFiltrace.ps1 -Action Restore -TargetRepository ../consumer -Configuration Release

The command requires PowerShell 5.1 or 7, Git, and the .NET 10 SDK. It does not need elevation, change a global tool installation, or upload the prepared package.

Using filtrace

Every analysis command takes a trace path and prints a dense text report (or compact JSON with --format json); collect instead launches the executable it records. The canonical investigation is orient -> rank -> drill -> compare: inspect the capture, rank the matching metric, drill an unwindowed CPU ranking when needed, then diff comparable traces or capture manifests against a baseline.

# Workflow: orient, rank the hottest frames, drill into one, then diff two runs.
filtrace info app.nettrace                     # 0. orient: format, providers, event counts, symbol rate
filtrace rank app.nettrace --metric cpu        # 1. what's hot (self weight: milliseconds or raw samples)
filtrace callers app.nettrace MyApp.Parse      # 2. who calls the hot frame
filtrace source app.nettrace --view lines --symbols bin/Release/net10.0   # 3. hot source lines
filtrace diff before.nettrace after.nettrace   # 4. what changed between runs
filtrace batch BenchmarkDotNet.Artifacts/filtrace-runs/run/manifest.json

The same analysis core is exposed as a stdio MCP server: every analysis command has a matching trace_* tool (eighteen in all - info -> trace_info, rank -> trace_rank, callers -> trace_callers, and so on), returning the same envelope shape and results. The capture and housekeeping commands (collect, cache) are CLI-only, as is widening an ETW analysis with --all-processes; MCP supports automatic or named-process scope. See Using filtrace from an AI agent for the client config and tool workflow.

Commands

Orient - see what a capture holds before ranking (the CLI counterpart of the trace_info tool):

Command Purpose Example
info Format, sample count, lost-event and symbol-quality warnings, supported analyses, and per-analysis capture/event state filtrace info app.nettrace

Ranking - rank stacks by a metric:

Command What it ranks Example
rank Any metric (cpu, alloc, exceptions, threadtime, contention, wait, activity) filtrace rank app.nettrace --metric contention

Implemented scope inventory:

  • Named process: CLI info, rank, source, callers, tree, classify, timeline, diff, batch, export, and report (for --kind diskio); MCP trace_info, trace_rank, trace_callers, trace_lines, trace_heatmap, trace_tree, trace_classify, trace_timeline, trace_diff, trace_batch, and trace_export, plus trace_diskio. Every listed surface except the disk report auto-scopes a multi-process .etl to the busiest process tree. Run processes / trace_processes first to inspect the capture, then set --process <name> / process to override. CLI commands expose --all-processes where an aggregate is supported; MCP trace_info and trace_rank expose allProcesses. Capture-wide info / trace_info preserves whole-capture totals but budget-limits its busiest-thread detail with an explicit truncation warning. Process inventory aggregates ETL and EventPipe CPU ownership without materializing frames; speedscope uses its full stack reader. Use info / trace_info for frame-name and source/PDB quality. Other stack-backed MCP analyses have no all-process aggregate; trace_diskio is the exception and, like the CLI disk report, remains machine-wide by default. When scoped, disk reports correlate DiskIOInit issuer IRPs to completions because a completion PID may be System or Idle. This is direct issuer scope, not causal ownership: deferred file-system/cache write-back issued by System is included only when System is selected, while Idle-issued I/O appears only in an unscoped report.
  • Exact process ids: the same commands and tools accept --pid <id>[,<id>] (comma-separated, not repeated) / pid instead of a name. A name substring is right for discovery, but a common host name such as dotnet matches every unrelated instance in a machine-wide capture and ranks them together; an exact id set cannot. Prefer it for manifests and automation. The three selectors are mutually exclusive, an id reused by two processes in one trace is refused rather than merged, and an id that is not in the trace is reported. For diff between two .etl traces with different process ids, supply both --before-pid <id>[,<id>] and --after-pid <id>[,<id>] (MCP: beforePid and afterPid). These scope the baseline and current traces separately, cannot be combined with the shared --pid, --process, or --all-processes selectors, and fail if either trace lacks a requested id. Manifest diffs instead use each case's recorded invocation ids; per-arm ids are not accepted for manifests.
  • Descendants: the same commands and tools accept --children include|exclude / children. Both selectors follow descendants by default, because the common capture shapes put the measured work in a child the host launched. Pass exclude to separate a parent's own CPU from a child runtime's; without it a native host's own cost is blended with the CoreCLR frames of the child it launched.
  • Disk completion time: CLI report --kind diskio and MCP trace_diskio accept --time <start>,<end> / time. The window filters completion timestamps after issuer-process IRP correlation; either bound may be open.
  • Invocation roots: CLI lifecycle and MCP trace_lifecycle take the same --process / --pid selectors, but each matched process instance is one invocation and descendants always follow, so neither takes --children or --all-processes.
  • Root subtree: CLI rank, callers, tree, classify, diff, batch, and export; MCP trace_rank, trace_callers, trace_tree, trace_classify, trace_diff, trace_batch, and trace_export. Set --root <frame> / root to keep the subtree under a frame. Root filtering is stack ancestry, not causal correlation: stacks without the selected frame are excluded, including sibling workers. Root-aware structured results identify rootKind: stackAncestry and report available versus retained weight and record counts; direct diffs report both sides, and manifest batch/diff report each case. Use an instrumented activity or validated time window for a parallel phase, and ETW threadtime when sampled CPU does not explain elapsed time.
  • BenchmarkDotNet workload: CLI rank, callers, tree, classify, diff, batch, and export accept --benchmark; MCP trace_rank, trace_callers, trace_tree, trace_classify, trace_diff, trace_batch, and trace_export accept benchmark: true. The preset isolates the WorkloadAction subtree from harness and overhead scaffolding; it is mutually exclusive with an explicit root. source views are not root-aware, so narrow them by method/file and treat percentages as process-scoped whole-trace values.

The rank command adds two more scopes: --activity <name> (the CPU samples taken inside one start-stop request/job) and --time <start>,<end> (milliseconds from the trace start, either bound optional; any metric on .nettrace / .etl), to zoom in on one request or the slice around a latency spike. Speedscope input is aggregate-only for --time and warns that the window was ignored.

filtrace rank bdn.nettrace --metric cpu --benchmark    # just the [Benchmark] code
filtrace rank bdn.nettrace --metric alloc --benchmark  # allocations under the workload
filtrace processes machinewide.etl             # list every process by weight
filtrace rank machinewide.etl --metric cpu --process MyApp   # one process tree
filtrace rank machinewide.etl --metric cpu --pid 9144,40356 --children exclude  # exactly those, parent-only
filtrace rank app.nettrace --time 1000,5000    # just the spike window

Native runtime symbols. Managed frames (including NGEN and ReadyToRun framework methods) resolve for free from the trace's CLR rundown. The unmanaged runtime frames - the GC, the JIT, memset / memcpy, write barriers - need PDBs from the Microsoft public symbol server, which rank fetches only when you opt in with --native-symbols (cached under --symbol-cache, default in the temp path). It is off by default so analysis stays offline and deterministic; the first run downloads, later runs hit the cache.

filtrace rank app.etl --metric cpu --process MyApp --native-symbols   # name the GC/JIT/memcpy frames

CPU drill-down - follow an unwindowed CPU ranking into detail:

Command Purpose Example
callers Immediate CPU callers of a frame, or a caller/callee view with --callees filtrace callers app.nettrace MyApp.Parse --callees
source Ranked source lines or per-file heat filtrace source app.nettrace --view heatmap --file Parser.cs
tree Top-down CPU call tree from the root filtrace tree app.nettrace --max-depth 5

Inventory - see what a (possibly machine-wide) capture contains:

Command Purpose Example
processes List processes by CPU-sample weight, to pick a --process target filtrace processes machinewide.etl
classify Summarize CPU weight by runtime work category (ms when established, otherwise samples) filtrace classify app.etl --native-symbols

Temporal - see what happened when, then scope a ranking to the busy window:

Command Purpose Example
timeline Aligned activity buckets or one bounded cross-lane snapshot filtrace timeline app.nettrace --mode snapshot --at 1500

Compare and export:

Command Purpose Example
diff Absolute/normalized CPU changes for traces or paired manifests filtrace diff before.nettrace after.nettrace
batch One compact ranking across every capture-manifest case filtrace batch run/manifest.json
export Write a flame graph (speedscope / chromium) filtrace export app.nettrace --format speedscope -o app.json

Structured reports:

Command Purpose Example
report GC, JIT, thread-pool, or physical disk-I/O report filtrace report app.nettrace --kind gc
lifecycle Per-invocation wall-clock phases: root lifetime, first child, child span, teardown (ETW) filtrace lifecycle run.etl --process myapp --image hostfxr
events Query raw events, filtered by name / payload / pid / tid, paged filtrace events app.etl --payload ConnectionReset

Capture (Windows, elevated) - record an ETW .etl yourself, no external recorder:

Command Purpose Example
collect Launch an executable and record a CPU, thread-time, startup, or physical disk-I/O .etl filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl --profile threadtime
filtrace collect --launch bin/Release/net10.0/MyApp.exe --output myapp.etl              # CPU
filtrace collect --launch dotnet --launch-args MyApp.dll --output tt.etl --profile threadtime
filtrace collect --launch MyApp.exe --output start.etl --profile startup                # low perturbation
filtrace collect --launch MyApp.exe --output io.etl --profile diskio                    # physical disk/files
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --working-directory C:\src\app --output build.etl
filtrace collect --launch cmd.exe --launch-args "/d /c build.cmd" --output build.etl --rundown --rundown-pid 1234 # known persistent CLR server
filtrace collect --launch MyApp.exe --output ring.etl --max-size-mb 512                 # bounded ring buffer

--working-directory sets and records the absolute directory inherited by every subject launch. Omit it to inherit the collector's current directory.

--rundown appends and merges a separate minimal CLR naming rundown after the launched command exits. Use it only when captured CPU belongs to managed servers that were already running and remain alive, such as compiler/build servers. The opt-in pass is machine-wide by default; pass up to 8 comma-separated exact ids to --rundown-pid to filter CLR naming events when the target servers are already known. The capture result records that filter. Rundown requests 512 MB of ETW buffers, and TraceEvent's derived maximum-buffer count permits the pool to grow to roughly 641 MiB. Provider activation can take up to 30 seconds, followed by up to 30 seconds of quiescence polling; the subsequent ETL merge has no timeout. Rundown can add hundreds of megabytes. It is not valid with --profile diskio or --max-size-mb; inspect the capture and info lost-event warnings before trusting resolved names. A targeted process must be running before capture and remain the same process through rundown; exit or PID reuse fails explicitly.

The diskio profile enables only physical DiskIO/DiskIOInit events, process/thread attribution, and the DiskFileIO name rundown. It deliberately omits CPU sampling, stacks, verbose FileIO, and CLR events. The capture and unscoped disk report are machine-wide; write the trace to a different volume when recorder writes must not contend with the workload volume. Scope a report to direct issuers with --process or --pid, and to completion time with --time.

With --format json, stdout contains only the capture-result JSON; identified subject stdout and stderr are forwarded to stderr. The command still exits successfully when capture succeeds even if the subject fails, and reports that failure in processExitCode and invocations. Text output continues to inherit the subject's streams.

For an EventPipe (.nettrace) capture - cross-platform, no elevation - use the first-party dotnet-trace (dotnet tool install -g dotnet-trace, then dotnet-trace collect -- <app>); collect is ETW-only.

File ops - manage the ETLX conversion cache filtrace keeps beside a trace:

Command Purpose Example
cache Build/reuse or remove the ETLX cache filtrace cache app.nettrace --action convert

ETLX conversion is coordinated per canonical trace path across threads and processes, with unique temporary files and atomic publication. Filtrace records the converter epoch, backend binary and source/cache file facts in <ETLX path>.filtrace.json (distinct from capture metadata at <trace path>.filtrace.json); a cache hit checks these small facts, not a hash of the entire trace. A readable ETLX with no valid marker is not reused or overwritten: run filtrace cache app.nettrace --action clean to explicitly discard it, then convert or analyze again. This one-time rebuild can take seconds for large traces. Owned caches rebuild when the source changes, and interrupted publications recover. clean removes both ETLX and its marker. Same-trace MCP queries may run in parallel; trace_info.etlxCacheState and cache --action convert report hit, waited, converted, or recovered.

Preview alias migration

The previous command names remain callable during the current migration window and print their canonical replacement to stderr, but they are hidden from top-level help and are not used in examples or generated guidance. Removal requires the explicit VN5 migration policy; it is not tied automatically to the passage of one preview release:

Previous names Canonical command
cpu, alloc, exceptions, threadtime rank --metric <name>
lines, heatmap source --view <name>
gcstats, jitstats, threadpool, diskio `report --kind gc
convert, clean `cache --action convert

Run filtrace <command> --help for the full option set of any command.

Using filtrace from an AI agent

filtrace is built for an agent mid-investigation. Two ways to wire it in:

  • MCP server - add the stdio server so the agent calls the trace_* tools directly:

    {
      "servers": {
        "filtrace": {
          "type": "stdio",
          "command": "dnx",
          "args": ["KlutzyNinja.Filtrace.Mcp", "--yes"]
        }
      }
    }
  • CLI - install the global tool (dotnet tool install -g KlutzyNinja.Filtrace) and let the agent shell out to filtrace <verb>.

Either way, the canonical loop is orient -> rank -> drill -> compare: read trace_info (CLI: filtrace info) first; when symbol resolution is below 0.8, inspect its warning and unresolved rows. Treat that as frame-name quality; before source-line analysis, inspect sourceResolution for exact matching PDB modules, mapped sampled managed frames, searched directories, highest-unmapped modules, and pdbIdentityMismatchModules. A mismatch means a same-named local PDB was found but its GUID or age differs from the trace. Once the relevant module matches, use sourceMappedManagedMethodCount versus sampledManagedMethodCount to confirm sampled methods resolve sequence points; unmappedNamedManagedFrameCount and highestUnmappedMethods expose the remaining <no source> impact. Use the generated BenchmarkDotNet child output when the outer build PDB does not match; use native symbols for CPU ETW runtime frames as applicable. Rank by the metric that matches the question (cpu, alloc, exceptions, threadtime, contention, wait, activity); for an unwindowed CPU ranking, drill the hot frame with callers / lines / tree; diff comparable CPU traces against a baseline.

Layout

Path Purpose
src/Filtrace.Core/ Analysis core: trace readers, stack-source providers, the provider-agnostic question-service engine. The only place logic lives.
src/Filtrace/ CLI host, packaged as the filtrace .NET global tool.
src/Filtrace.Mcp/ Stdio MCP host, packaged separately for dnx KlutzyNinja.Filtrace.Mcp.
benchmarks/Filtrace.Benchmarks/ BenchmarkDotNet performance harness for the analysis core.
benchmarks/Filtrace.PerfWorkload/ Parameterized CPU/activity workload for reproducible Track D traces.
tests/Filtrace.Core.Tests/ Unit + golden-file contract tests.
tests/Filtrace.Parity.Tests/ Numeric parity against the frozen legacy oracles.
eval/ Headless-agent eval harness, tasks, baselines.
docs/ Design, roadmap, competitive analysis, and the single-source workflow text for the skill / README / help.
.agents/skills/filtrace/ The shipped agent skill.

Self-containment

filtrace carries its own Directory.Build.props, Directory.Build.targets, Directory.Packages.props, global.json, and .editorconfig (root = true), so the build is fully self-contained. Its only external dependency is the published KlutzyNinja.Touki NuGet package; it references no other project.

Build and test (standalone)

cd filtrace
dotnet build filtrace.slnx
dotnet test filtrace.slnx

About

Command-line and MCP .NET trace analyzer: rank, drill, diff, and export CPU / allocation / exception / thread-time profiles from .nettrace, .etl, and speedscope captures.

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages