Site Reliability Engineer II
Owned production reliability for Kubernetes microservices on AKS, GitLab CI/CD pipelines, and Linux-based operational automation — and built an LLM-powered incident-triage service that turns logs and telemetry into actionable recommendations for on-call engineers.
- CI/CD automation
- Kubernetes reliability
- LLM-assisted operations
- Release engineering
- Production support
- Linux automation
- On-call RCA
- Built a Python/FastAPI incident-triage service that feeds Kubernetes logs, deployment configuration, and incident context to an LLM endpoint to generate troubleshooting recommendations for on-call engineers.
- Engineered LLM output guardrails: Pydantic schemas validate structured responses before they reach engineers, and recommendations are checked against live telemetry before being codified as runbook procedures.
- Owned production reliability for Kubernetes microservices on AKS using Helm; tuned CPU/memory requests and limits, liveness/readiness probes, HPA autoscaling, deployment strategies, pod scheduling, and ephemeral-storage limits to improve stability and reduce infrastructure waste.
- Owned on-call root-cause analysis for production AKS incidents, tracing failures across pod, node, storage, and network layers — including CrashLoopBackOff, ImagePullBackOff, OOMKilled pods, failed readiness probes, storage pressure, node resource exhaustion, and service discovery — using kubectl, Grafana, journalctl, tcpdump, and Linux performance commands.
- Built GitLab CI/CD pipelines using Python, Bash, YAML, and Docker with secrets handling, approval gates, smoke tests, pre-deployment checks, rollback stages, and environment promotion to reduce failed deployments and production impact.
- Automated Linux and Kubernetes health checks, remediation, patching, environment provisioning, and configuration templates using Python, Bash, Ansible, systemd, and cron to reduce drift and manual toil.
- Diagnosed NGINX/Ingress routing, DNS resolution, TLS certificates, and Azure networking issues affecting service connectivity.
- Collaborated with development, platform, security, and cloud infrastructure teams to troubleshoot CI/CD failures, Kubernetes platform issues, networking dependencies, access/RBAC problems, and service reliability risks.
- AKS
- Kubernetes
- Helm
- GitLab CI/CD
- Docker
- NGINX / Ingress
- Azure Networking
- DNS
- TLS
- Linux
- Ansible
- systemd
- Bash
- Python
- FastAPI
- Pydantic / LLM Guardrails
- HPA / Autoscaling
- Grafana
- Prometheus
- kubectl
- RBAC
- On-call / Incident Response