I run production infrastructure for game and product companies as a contractor — AWS/EKS, Terraform, CI/CD, observability, and AI-agent automation for ops. Fixed-scope projects and monthly retainers.
Legacy ELK, scattered logs, no tracing — root-cause analysis took hours. Migrated logging to Grafana Loki (S3-backed), added Tempo tracing with logs↔traces correlation, built alerting, uptime dashboards and on-call runbooks.
→ MTTR from hours to minutes; log storage on cheap S3 instead of an Elasticsearch cluster.
A web game with live traffic needed deploys that don't break sessions already open in players' browsers. Built the platform on AWS EKS + CloudFront, fully in Terraform, with GitLab CI/CD and a three-layer client caching model (immutable assets / no-cache entry points / localStorage).
→ Deploys during live traffic, zero downtime, no broken client sessions.
Moved a product's CDN between providers with no user-visible breakage, then tracked down a subtle post-migration bug: conditional CORS headers on the new edge broke Next.js client-side navigation for cached responses. Fixed at the edge configuration level.
→ No downtime; eliminated the "works on refresh, breaks on navigation" error class.
Built agentic pipelines on top of Claude: multi-agent workflows with structured outputs, verification gates and human-in-the-loop approval for irreversible steps. Agents do the investigation — querying live systems, correlating logs and configs; humans only approve state changes. Full audit trail.
→ Hours of routine ops per week moved off senior engineers.
A defined result with a deadline and a price: observability stack, platform build-out, CDN/cloud migration, cost audit. Typically 2–4 weeks.
Your infrastructure stays healthy: incidents handled, releases shipped, monitoring watched. A senior engineer on call without a full-time hire.
Your DevOps seat is open? I cover the gap from week one and hand everything over, documented, when you hire.
Cloud: AWS (EKS, CloudFront, S3, IAM), Yandex Cloud ·
IaC: Terraform, Helm ·
CI/CD: GitLab CI, GitHub Actions, ArgoCD
Observability: Grafana, Loki, Tempo, Prometheus ·
Runtime: Kubernetes, Docker ·
Automation: Python, Bash, Claude agent pipelines