Skip to content

Latest commit

 

History

42 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Smart Network Troubleshooting Agent

An autonomous AI agent that diagnoses network failures like a Senior Network Engineer and explains them in plain English — so non-technical users don't need to know what ping or traceroute even means.

Python License Status CS5700


The Problem

When a home network breaks, most users are completely lost.

Traditional tools like ping, traceroute, and nslookup generate raw technical output that requires networking expertise to interpret. Worse, users don't even know which tool to run first — or in what order.

$ ping google.com
Request timeout for icmp_seq 0
Request timeout for icmp_seq 1
^C

--- google.com ping statistics ---
2 packets transmitted, 0 packets received, 100.0% packet loss

What does this mean? Is it DNS? Routing? A local adapter? For 90% of users, this is an alien language.

Existing solutions don't help:

  • Raw CLI tools (iputils, RIPE Atlas) — powerful, but require expert interpretation
  • Consumer troubleshooters — just say "check your connection"
  • AI log analysis (LogPAI, OpenTelemetry) — designed for backend operators, not end users

The gap: there are plenty of tools that generate network data. What's missing is an intelligent layer that reads that data and tells a non-technical user exactly what's wrong and what to do.


Our Solution

A diagnostic agent that:

  1. Accepts a natural language symptom: "I can't load any websites"
  2. Automatically selects and runs the right diagnostic tools
  3. Reasons about the results to identify the root cause
  4. Returns a plain-English explanation with actionable recommendations

No terminal knowledge required. No manual tool selection. No raw output to interpret.


How It Works

Iter 1 — Foundation (Complete)

User: "I can't load any websites"
         │
         ▼
   CLI or Web Interface
         │
         ▼
   extract_target()          →  finds "8.8.8.8" or a domain from the symptom
         │
         ▼
   run all 3 tools (fixed order)
   PingTool → DNSTool → TracerouteTool
         │
         ▼
   assemble results into structured message
         │
         ▼
   LLM API call (OpenAI / Anthropic / Google / Groq)
         │
         ▼
   structured JSON response
   { summary, root_cause, recommendations }
         │
         ▼
   plain-English output to user

Iter 2 — ReAct Loop (Complete)

Instead of running all tools blindly, the agent uses a ReAct (Reasoning + Acting) loop to make decisions between each tool call:

Think: "General connectivity issue — check baseline first"
Act:   run PingTool
Observe: 0% packet loss — network layer is fine

Think: "Network OK, issue might be DNS"
Act:   run DNSTool
Observe: DNS resolution failing — NXDOMAIN

Think: "Root cause identified. No need for traceroute."
Stop → deliver diagnosis

This mirrors how a human network engineer actually troubleshoots — adaptively, based on evidence — rather than running every tool every time.

Iter 3 — Evaluation Framework (Complete)

Added a fourth tool (curl) for HTTP-layer diagnosis, a Docker sandbox with 10 fault injection scenarios, and an automated evaluation framework.

Result: 90% diagnostic accuracy on ambiguous user descriptions vs 50% for a naive LLM baseline (gemini-2.5-flash).


Features & User Stories

Feature 1 — Symptom-Based Automated Diagnosis ✅ (Iter 1)

Who: As a non-expert home user experiencing connectivity problems

Why: So that I can understand what is wrong and fix it without networking knowledge

What: I want to describe my problem in plain English and receive a clear diagnosis with specific next steps

Acceptance criteria:

  • User inputs symptoms in natural language via CLI or web interface
  • System automatically runs appropriate diagnostic commands
  • LLM output is in plain English — no raw command output shown
  • At least 2 specific, actionable recommendations provided

Feature 2 — Intelligent Tool Selection with Visible Reasoning ✅ (Iter 2)

Who: As a networking student learning how to troubleshoot

Why: So that I can observe how an expert systematically narrows down problems

What: I want the agent to show its reasoning before each tool execution

Acceptance criteria:

  • Agent displays "Thought:" reasoning before each tool run
  • Tool selection adapts based on previous results — not a fixed sequence
  • Different symptoms lead to different diagnostic paths
  • Agent stops when it has enough evidence, not always running all 3 tools

Feature 3 — Reproducible Fault Injection Testing Environment ✅ (Iter 3)

Who: As a developer or instructor evaluating the agent

Why: So that I can validate the agent's diagnostic accuracy objectively

What: I want a controlled environment to inject specific faults and measure how accurately the agent diagnoses them

Acceptance criteria:

  • Docker environment starts with docker compose up
  • 10 fault scenarios injectable (DNS failure, packet loss, high latency, etc.)
  • Accuracy measured by predicted root_cause == ground truth label — not keyword matching
  • Benchmark produces 90% agent accuracy vs 50% naive LLM baseline on ambiguous symptom descriptions

Supported LLM Providers

Model Provider Cost/run Free Tier Notes
gpt-4o-mini OpenAI ~$0.0005 ❌ Default — best value
gpt-4o OpenAI ~$0.008 ❌ Final evaluation
claude-haiku-4-5-20251001 Anthropic ~$0.001 ❌ Agent-task optimized
gemini-2.5-flash Google ~$0.0002 ✅ Cheapest capable model
llama-3.3-70b-versatile Groq ~$0.0004 ✅ Open source, data stays local

All providers use the same interface — swap models with a single --model flag.


Setup

# Clone
git clone https://github.com/Zeglow/network-diagnostic-agent.git
cd network-diagnostic-agent

# Install dependencies
pip install -r requirements.txt

# Configure API keys
cp .env.example .env
# Edit .env — add the keys for whichever providers you want to use

.env format:

OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
GOOGLE_API_KEY=...
GROQ_API_KEY=gsk_...

Usage

CLI

# Default model (gpt-4o-mini)
python src/cli.py diagnose "I can't load any websites"

# Specific model
python src/cli.py diagnose "My video calls keep dropping" --model gemini-2.5-flash

# List all available models
python src/cli.py diagnose

Example output:

Analyzing: I can't load any websites
Model: gpt-4o-mini
Starting ReAct loop...

  Step 1: asking LLM what to do next...
  Thought: General connectivity issue — check baseline first.
  Action:  run ping
  Observation: OBSERVATION: ping result:
Status: SUCCESS
Data: {'packet_loss_percent': 0.0, 'avg_rtt_ms': 12.3}...

  Step 2: asking LLM what to do next...
  Thought: Network layer is fine. The issue might be DNS.
  Action:  run dns
  Observation: OBSERVATION: dns result:
Status: FAILED
Error: DNS lookup failed for 8.8.8.8...

  Step 3: asking LLM what to do next...
  Thought: DNS resolution is failing. I have enough to diagnose.
  → Diagnosis ready after 3 step(s), 2 tool(s) used

## Reasoning Trace

Step 1:
  Thought: General connectivity issue — check baseline first.
  Action:  run ping

Step 2:
  Thought: Network layer is fine. The issue might be DNS.
  Action:  run dns

Step 3:
  Thought: DNS resolution is failing. I have enough to diagnose.
  Action:  provide final diagnosis

(3 step(s) taken, 2 tool(s) used: ping, dns)

## Summary
Ping to 8.8.8.8 is succeeding but DNS resolution is failing,
suggesting your DNS server is unreachable.

## Root Cause
DNS failure — your device cannot resolve hostnames to IP addresses.

## Recommendations
1. Flush your DNS cache: sudo dscacheutil -flushcache (macOS)
2. Switch your DNS server to 8.8.8.8 in your network settings
3. Restart your router if the issue persists

Web Interface

python app.py
# Open http://localhost:5001

The web interface includes a model selector — no terminal knowledge required.


Project Status

Iteration Weeks Feature Status
Iter 1 5–7 Basic diagnostics + LLM explanation + web UI ✅ Complete
Iter 2 8–9 ReAct loop — adaptive tool selection ✅ Complete
Iter 3 10–11 Docker fault injection + evaluation + curl tool ✅ Complete
Final 12–13 Report + presentation ⬜ Planned

Evaluation Strategy (Iter 3)

We evaluate against 10 Docker-injected fault scenarios with ground truth labels:

# Scenario Injected Fault Ground Truth Label
1 DNS Failure Replace /etc/resolv.conf with invalid DNS dns_failure
2 High Packet Loss tc netem loss 30% packet_loss
3 High Latency tc netem delay 500ms high_latency
4 Route Failure ip route add blackhole route_failure
5 Port Blocked iptables REJECT port 80 port_blocked
6 Complete Outage Block all egress no_connectivity
7 Intermittent Loss tc netem loss 10% 25% intermittent_loss
8 Bandwidth Throttle tc tbf rate 100kbit bandwidth_throttle
9 High Jitter tc netem delay 100ms 80ms high_jitter
10 Duplicate Packets tc netem duplicate 20% duplicate_packets

Accuracy is measured by comparing predicted["root_cause"] == ground_truth_label — not keyword matching.

ReAct agent achieves 90% accuracy vs 50% naive LLM baseline on ambiguous symptom descriptions (gemini-2.5-flash).


Why This Is Not Just a Wrapper

Most "AI + networking" projects simply pass raw tool output to ChatGPT and display the response. This project differs in three ways:

1. Adaptive reasoning (Iter 2): The ReAct loop makes decisions between tool calls based on accumulated evidence — the same pattern used in production systems like GitHub Copilot Workspace and AWS Q Developer.

2. Rigorous evaluation: Structured JSON output with root_cause enum enables automated accuracy measurement against ground truth labels. This is how production AI systems are evaluated, not how demos are built.

3. Multi-provider, open-source support: The same agent runs across 5 LLM providers including local open-source models (Llama via Groq). No data needs to leave your machine.


Team

Ashley (Yongqi) Ou — Agent core, LLM integration (multi-provider), prompt engineering, curl tool, web interface, evaluation framework, project coordination

Avery (Weiyu) Qiu — Diagnostic tool wrappers, Docker sandbox, fault injection scripts, CLI interface, demo videos

CS 5700 Fundamentals of Computer Networking — Northeastern University, Spring 2026

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages