An ML-based system for automatically detecting flood peaks from hydrograph data using manually labeled peaks for training.
PeakPicker uses machine learning to automate the identification of flood peaks in river gage streamflow data. The system is trained on manually selected peaks and can:
- Train models on gages with known return periods
- Detect peaks in existing gage data
- Process new gage files and generate predictions
- Calculate return periods when not available
- Generate visualizations of hydrographs with detected peaks
This project originated during research for the paper "How Well Do U.S. National Water Model Short-Range Forecasts Predict Flood Event Timing and Magnitude?" While analyzing flood events for selected gages, we needed to optimally identify peaks in hydrographs that corresponded to actual flood events with specific characteristics. The manual labeling process for 306 gages proved to be time-intensive and challenging.
To streamline this process and enable analysis of larger datasets in future studies, we developed PeakPicker—a machine learning model that automates the identification of event-related peaks in streamflow hydrographs. What began as an internal tool has evolved into a robust system that can assist researchers and practitioners in efficiently detecting flood peaks across extensive gage networks.
Choose your use case:
If you want to see how the model works, train it, and test on sample data:
-
Follow the Quick Start section below to:
- Train the model on included sample data
- Test peak detection on existing gages
- See example outputs and visualizations
-
Read the QUICKSTART.md for step-by-step tutorial
-
Check INSTALLATION.md for installation help
If you have your own gage data to process in a custom project folder:
-
Use the project runner:
project/run_project.py -
Read project/PROJECT_RUNNER_GUIDE.md for complete instructions on:
- Setting up your project folder structure
- Preparing your data files (gages, return periods, mappings)
- Running batch processing on multiple gages
- Customizing output locations
-
Quick example:
# Process all gages in your project folder python project/run_project.py --project-dir /path/to/your/project # Process a single gage python project/run_project.py --gage 01200600
Key Difference: The main peakpicker.py script uses the included sample data, while project/run_project.py is designed for processing your own custom datasets with your own folder structure.
peakpicker/
├── Core Package (for testing/development)
│ ├── data_loader.py # Data loading and preprocessing
│ ├── feature_engineering.py # Feature extraction from time series
│ ├── return_period_calculator.py # Return period calculations
│ ├── model_trainer.py # ML model training
│ ├── plotter.py # Visualization tools
│ ├── peakpicker.py # Main script for sample data
│ ├── requirements.txt # Python dependencies
│ ├── venv/ # Virtual environment
│ ├── gages/ # Full sample gage data files (300+ gages)
│ ├── gages_samples/ # 10 sample gages for quick testing
│ │ └── *_Obs.csv # Format: datetime_utc, discharge_cms
│ ├── manual_added_peaks.csv # Manually labeled peaks
│ ├── return_periods.csv # Return period values
│ └── plots/ # Generated plots (created automatically)
│
├── Production Runner (for your data)
│ └── project/ # Project folder with runner and data
│ ├── run_project.py # Project runner script
│ ├── PROJECT_RUNNER_GUIDE.md # Complete usage guide
│ ├── gages/ # Your gage CSV files
│ ├── data/ # Your return periods & mappings
│ ├── plots/ # Output plots (auto-created)
│ └── results/ # Output CSVs (auto-created)
│
└── Documentation
├── README.md # This file
├── QUICKSTART.md # Step-by-step tutorial
└── INSTALLATION.md # Installation & setup guide
# On macOS/Linux
source venv/bin/activate
# On Windows
venv\Scripts\activateThe virtual environment is already created with all necessary packages installed.
python -c "import pandas, sklearn, xgboost; print('All packages installed!')"Train a peak detection model on all gages with return periods:
python peakpicker.py --trainThis will:
- Load all gages with return period data
- Extract features from streamflow time series
- Train an XGBoost classifier
- Evaluate performance on test data
- Save the model to
peak_model.pkl
python peakpicker.py --gage 03408500This will:
- Load the gage data
- Apply the trained model
- Detect flood peaks
- Generate plots
- Save results to
03408500_detected_peaks.csv
For a new gage with format: datetime,streamflow
python peakpicker.py --file new_gage.csv --gage-number 12345678This will:
- Load the new file
- Calculate return periods from the data
- Detect peaks
- Generate plots
- Save results
# XGBoost (default, recommended)
python peakpicker.py --train --model-type xgboost
# LightGBM (faster training)
python peakpicker.py --train --model-type lightgbm
# Random Forest
python peakpicker.py --train --model-type random_forest# More sensitive (detects more peaks)
python peakpicker.py --gage 03408500 --threshold 0.3
# Less sensitive (detects fewer peaks)
python peakpicker.py --gage 03408500 --threshold 0.7# Peaks must be at least 72 hours apart
python peakpicker.py --gage 03408500 --min-distance 72
# Peaks must be at least 24 hours apart
python peakpicker.py --gage 03408500 --min-distance 24python peakpicker.py \
--file data.csv \
--gage-number 12345678 \
--datetime-col timestamp \
--discharge-col flow_ratepython peakpicker.py --gage 03408500 --no-plotpython peakpicker.py --gage 03408500 --output my_peaks.csvCSV files should contain:
- datetime column: Timestamps (any standard format)
- discharge column: Streamflow values (in cms or any consistent unit)
Example:
datetime,streamflow
2021-01-01 00:00:00,45.2
2021-01-01 00:15:00,45.5
2021-01-01 00:30:00,45.8Detected peaks are saved as CSV with columns:
peak_time_utc: Timestamp of detected peakpeak_flow_cms: Discharge value at peakpeak_probability: Model confidence (0-1)site_no: Gage number
Return periods represent the statistical recurrence interval of flood events:
- 2-year: Expected once every 2 years (common flood)
- 5-year: Expected once every 5 years
- 10-year: Expected once every 10 years
- 25-year: Expected once every 25 years
- 50-year: Expected once every 50 years
- 100-year: Expected once every 100 years (rare, extreme flood)
The system can:
- Use pre-calculated return periods from
return_periods.csv - Calculate return periods from streamflow data using statistical methods
The model uses 100+ features including:
Temporal Features:
- Hour, day, month, season
Statistical Features:
- Rolling means, standard deviations
- Percentiles (50th, 75th, 90th, 95th, 99th)
- Skewness and kurtosis
Derivative Features:
- Rate of change (velocity, acceleration)
- Sign changes
Local Features:
- Local maxima detection
- Distance to local peaks
- Peak prominence and width
Return Period Features:
- Ratios to return period thresholds
- Exceedance indicators
The system handles gage numbers with/without leading zeros:
03408500and3408500are treated as the same gage- Files are matched flexibly (e.g.,
03408500_Obs.csvor3408500_Obs.csv)
You can also use PeakPicker programmatically:
from peakpicker import PeakPicker
# Initialize
picker = PeakPicker(model_path='peak_model.pkl')
# Train model
picker.train_model(model_type='xgboost')
# Detect peaks for existing gage
detected_peaks = picker.detect_peaks_for_gage(
site_no='03408500',
probability_threshold=0.5,
min_peak_distance_hours=48,
plot=True
)
# Process new file
detected_peaks = picker.process_new_gage_file(
file_path='new_gage.csv',
gage_number='12345678',
datetime_col='datetime',
discharge_col='streamflow',
plot=True
)The system creates three types of plots:
-
Hydrograph with Peaks (
{gage}_hydrograph.png)- Time series of streamflow
- Detected peaks marked in red
- Manual peaks marked in green (if available)
- Return period thresholds
- Peak probability overlay
-
Peak Comparison (
{gage}_peak_comparison.png)- Comparison between detected and manual peaks
- Distribution analysis
- Monthly patterns
- Statistical summary
-
Annual Maxima (
{gage}_annual_maxima.png)- Annual maximum flows
- Return period thresholds
- Trend analysis
All plots are saved to the plots/ directory (created automatically).
The model is evaluated using:
- Precision: Proportion of detected peaks that are correct
- Recall: Proportion of actual peaks that were detected
- F1-Score: Harmonic mean of precision and recall
- AUC: Area under ROC curve
During training, you'll see:
Metrics:
precision: 0.8523
recall: 0.7891
f1_score: 0.8194
auc: 0.9234
FileNotFoundError: Model file not found: peak_model.pklSolution: Train the model first:
python peakpicker.py --trainWarning: Could not find data file for gage 12345678Solutions:
- Check that the gage file exists in the
gages/directory - Verify the file name format:
{gage_number}_Obs.csv - Try with/without leading zeros
If you encounter memory errors during training:
- Process fewer gages at once (modify
data_loader.py) - Reduce feature window sizes in
feature_engineering.py - Use a lighter model:
--model-type random_forest
# Test data loader
python data_loader.py
# Test feature engineering
python feature_engineering.py
# Test return period calculator
python return_period_calculator.py
# Test plotter
python plotter.py
# Test model trainer
python model_trainer.pyEdit model_trainer.py to adjust:
- Number of estimators
- Learning rate
- Max depth
- Class weights (for handling imbalance)
Edit feature_engineering.py to add new features:
def add_custom_features(self, df):
# Add your custom features here
df['my_feature'] = ...
return dfThen add to engineer_features() method.
gages/*_Obs.csv: Gage streamflow datamanual_added_peaks.csv: Manually labeled peaks for trainingreturn_periods.csv: Return period values (optional, can be calculated)
peak_model.pkl: Trained ML model{gage}_detected_peaks.csv: Detected peaks for each gageplots/*.png: Visualizations
If you have your own gage datasets to process, use the project runner instead of the main peakpicker.py script:
your_project/
├── gages/ # Your gage CSV files (datetime, discharge)
├── data/
│ ├── usgsid_comid.csv # Maps USGS IDs to COMIDs
│ └── return_periods.csv # Return periods by COMID
├── plots/ # Output plots (auto-created)
└── results/ # Output CSVs (auto-created)
# Process all gages
python project/run_project.py --project-dir /path/to/your_project
# Process single gage
python project/run_project.py --project-dir /path/to/your_project --gage 01200600
# Adjust sensitivity
python project/run_project.py --project-dir /path/to/your_project --threshold 0.6See project/PROJECT_RUNNER_GUIDE.md for:
- Detailed setup instructions
- Data file format specifications
- Batch processing examples
- Troubleshooting tips
- Customization options
Key Features:
- ✅ Batch process multiple gages
- ✅ Custom folder structure
- ✅ Automatic return period matching
- ✅ Handles USGS IDs with/without leading zeros
- ✅ Generates combined results CSV
- ✅ Beautiful visualizations with colored return period zones
If you use this tool in your research, please cite:
PeakPicker - Automated Flood Peak Detection
Brigham Young University
2025
For issues or questions:
- Check this README for solutions
- Review the example outputs
- Examine the generated plots for insights
[Specify your license here]
This project was developed as part of research at Brigham Young University for automated flood peak detection using machine learning on hydrologic time series data.
Note: This repository represents a side result from a research paper and project conducted at BYU. Upon completion of necessary approvals, this repository will be transferred to the official BYU Hydroinformatics organization.
Developed at Brigham Young University's Hydroinformatics Lab for advancing automated hydrologic analysis and flood detection methodologies.