SRE who has managed 50+ Kubernetes clusters at 99% uptime across AWS, Azure, and GCP — cut MTTR by 40% through observability automation, and built LLM-powered AIOps tools that resolve incidents before users notice.
Production microservices on EKS with Terraform, ArgoCD, Prometheus, Grafana, and AWS Bedrock+Lambda AIOps assistant for natural language querying of live metrics.
NL querying of live production metrics via Bedrock — real-time AIOps without manual dashboard lookup
EKSTerraformArgoCDAWS BedrockLambda
Slurm · GitLab CI · AIOps
Slurm ML Platform
Production AI research cluster: Slurm/Docker, FastAPI job management, AIOps anomaly monitor, GPU health validation, GitLab CI/CD with 42 automated tests.
42-test CI/CD suite with GPU health validation — production-grade AI research infrastructure
SlurmDockerFastAPIKubernetesGitLab CI
Claude API · RAG · ServiceNow
LLM SRE Runbook Generator
Claude API + RAG over postmortem docs. Terraform-managed Lambda webhook auto-delivers AI runbooks into ServiceNow incidents on Prometheus alert trigger.
3-hour manual runbook creation → under 10 minutes — auto-delivered to ServiceNow on alert fire
Claude APIRAGLambdaTerraformServiceNow
LLM-Benchmark Suite
Real-time LLM inference benchmarking: streams 4 providers (Groq, Cerebras, Cloudflare, Fireworks) in parallel via SSE with TTFT, throughput, P95 latency, and cost metrics.
Node.jsExpressSSEGroqCerebras
RUN:ai GPU Scheduler
Self-hosted GPU job scheduling on k3s NVIDIA A10 VM — FastAPI + Next.js + Redis/RQ via Helm, DCGM exporter for GPU metrics, SSE live log streaming.
k3sNVIDIAFastAPINext.jsHelm
LLM Deployment on GPU Cluster
End-to-end LLM inference stack on EKS g4dn spot with vLLM (Llama-3.1-8B), DCGM → Prometheus → Grafana — all infra as Terraform IaC.
EKSTerraformvLLMDCGMPrometheus
LLM-GitOps-Pipeline
MLOps pipeline: MLflow model versioning, FastAPI/Docker inference server, ArgoCD GitOps canary releases via nginx, Prometheus/Grafana observability.
vLLM + Triton on k3s NVIDIA VM, KEDA event-driven autoscaling triggered by Prometheus queue-depth, benchmarked 100 concurrent users with Locust.
vLLMTritonk3sKEDALocust
Serverless vs Self-Hosted LLM
Benchmarked Mistral-7B on Modal A10G vs EKS/T4: quantified 10x cold-start gap, 3x TCO difference, ~15 req/min break-even for enterprise GPU platform selection.
ModalEKSMistral-7BPythonAWS
FineTune
Full-stack LLM fine-tuning platform with LoRA training on Modal GPU, HuggingFace Hub auto-push, and real-time log streaming via Supabase.
Next.jsFastAPIModal GPULoRAHuggingFace
Gemini Multi-Agent Support Platform
Multi-agent LLM pipeline on GCP (LangGraph + Gemini) with triage/specialist/escalation routing, FAISS/Vertex AI RAG, streaming UI, Cloud Run via Terraform.
Self-reflecting GenAI pipeline on Google ADK: parallel specialist agents (RAG/web/cost) feeding a synthesis agent that challenges its own output; Cloud Run auto-deploy.
Google ADKGeminiCloud RunGCP
LLM Observability & Cost Intelligence
Production observability for multi-agent LLMs: per-call token counts, latency histograms, USD cost attribution across providers, anomaly alerting at 3× rolling baseline.
AnthropicOpenAIGeminiPrometheusPython
Claude Deployment Eval Framework
Full-stack LLM eval platform (FastAPI + Next.js + PostgreSQL): auto-generates adversarial test suites for Claude deployments, streams diagnostics via SSE, clusters failures with LLM-as-judge.
FastAPINext.jsPostgreSQLClaude APISSE
ServiceNow MCP Server
Production MCP server connecting Claude to ServiceNow — 9 tools for incident triage, KB search, CMDB dependency mapping; 87% P1/P2 priority accuracy on 35-case eval suite.
Secure MCP gateway with local DuoGuard-0.5B classifiers for inbound prompt injection screening and real-time outbound PII masking (GDPR/HIPAA compliant).
FastAPIMongoDBGeminiPython
PerceptionSentinel
Fault-tolerant AV perception: YOLOv8n primary + LLM fallback (NVIDIA NIM) validated by NeMo Guardrails; SRE watchdog redirects vehicle control on SLO breach.
YOLOv8NVIDIA NIMNeMoPrometheusPython
TemporalOps
Self-healing K8s canary orchestrator in Go on Temporal: saga-based auto-rollback, human approval gate, crash-resilient execution, Kyverno policies, audit trail.
GoTemporalKubernetesKyvernoPrometheus
SentryCTL
Linux incident debugger CLI in Go — simulates and diagnoses production faults via stress-ng, strace, perf, tcpdump, and eBPF.
GoeBPFperfstraceLinux
CapSim
Capacity-planning and autoscaling simulator on Kubernetes: Erlang-C (M/M/c) queueing model with regression traffic forecasting predicts replicas 30s ahead and preemptively raises HPA floors — full capacity within ~10s of a 6× spike while reactive CPU-based HPA sat 4× under-provisioned.
GoKubernetesk6PrometheusGrafana
MySQL HA
HA MySQL cluster (primary + 2 replicas, GTID semi-sync) with automated failover via ProxySQL and a custom Python controller — chaos-validated (process kill, network partition, disk exhaustion) with sub-9s MTTR and zero read-path downtime.
MySQLProxySQLPythonChaos Engineering
ChipFarm
Containerized HPC cluster (SLURM, LDAP/SSSD, NFS/autofs, Prometheus/Grafana) with LLM-assisted triage CLI — detects stuck jobs and drained nodes, proposes remediations via NVIDIA NIM, executes only with human approval.
SLURMNVIDIA NIMPrometheusLDAPNFS
experience
NestleSoftware Engineer · ContractJan 2026 – Present · Seattle, WA
ArgoCDAzure DevOpsACR
SAP America IncDevOps Engineer · InternJul 2024 – Sep 2025 · Chicago, IL
KubernetesGoOpenTelemetry
Northeastern UniversityGraduate TA · Cloud Computing & RDBMSSep 2023 – Apr 2024 · Boston, MA
AWSKubernetesTerraform
United Online IncSoftware Quality EngineerDec 2020 – Jul 2022 · Hyderabad, IN