Building high-scale real-time platforms, distributed AI inference clusters, and resilient cloud architectures.
Staff Software Engineer and Founding Engineer with 5+ years of experience architecting high-throughput distributed platforms, real-time media systems, and production AI infrastructure. Driven by a "Systems First" engineering philosophy — prioritizing sub-10ms latencies, high availability, zero-trust observability, and rapid developer velocity.
- Infrastructure at Scale: Architected zero-to-one platform infrastructure for Vooz Inc., scaling to 500,000+ Monthly Active Users (MAU) with zero downtime.
- Distributed AI Inference: Migrated standalone ML models to an NVIDIA Triton Inference Server cluster featuring dynamic batching and GPU optimization, scaling throughput to 1,280+ RPS.
- Edge Intelligence: Engineered zero-latency, client-side browser inference using ONNX Runtime (WebGPU / WASM) and Cache API for offline-capable AI features.
- Real-Time Financial Settlement: Designed Solana microservices integrated with Helius Webhooks for atomic payment confirmations and real-time point balance hydration.
- Resilient Media Recovery: Built automated WebRTC connection recovery mechanisms and state synchronization for fault-tolerant audio/video streaming.
- Systems-First Resilience: Design for explicit failure modes, distributed tracing, and graceful degradation under extreme load.
- Edge ML Acceleration: Offload inference compute to browser-local WebGPU/WASM models to eliminate network round-trips.
- Immutable Infrastructure: Enforce GitOps continuous deployment, infrastructure as code, and declarative cluster management.
- Data-Driven Performance Tuning: Rely on OpenTelemetry profiling and metrics to resolve micro-bottlenecks before scaling horizontal compute.



