Senior DevOps / Platform / SRE Engineer · AI Infrastructure & LLMOps

Anirudh Vaka

Senior DevOps / Platform engineer — production infra on AWS, Azure & Kubernetes at 99.9% uptime for 1000+ customers, plus self-hosted LLMs and AI-in-SDLC auto-remediation. Founder of two live AI SaaS products.

  • Leads a team of 5
  • Intern → DevOps Lead in < 2 yrs
  • 200+ CI/CD pipelines
  • ISO 27001:2022

Serving notice — available from 23 Oct 2026 · open to visa sponsorship / relocation · fully remote-friendly

I build and operate production infrastructure — and I ship products on top of it. By day, I lead DevOps for an enterprise SaaS platform and operate an on-prem Kubernetes data center I built from bare metal. Outside work, I run two paid SaaS products — PrepAtlas (AI-assisted exam prep) and HumanifyCV (AI text humanization). Currently exploring AI Infrastructure / MLOps, Platform, and Senior DevOps / SRE roles — open to relocation worldwide with visa sponsorship, or fully remote.

Uptime
0.0%
on-prem K8s, 2 years
Pipelines
0+
GitHub Actions + Azure DevOps
Faster releases
0%
via containerization
Cloud cost cut
0%
FinOps + right-sizing
Paid SaaS products
0
PrepAtlas, HumanifyCV
Production DevOps
0+ yrs
since Jan 2023

Selected Projects

Two products I run end-to-end with paying users, the AI platform I built for internal engineering, and the production infrastructure I architect and operate for work.

Product · Live

PrepAtlas

AI learning + career platform — tutor, adaptive practice, exam engine, role tracks

The problem. Exam prep is sold as content — a fixed syllabus, recorded lectures, one generic question bank. None of it knows which topic a particular student is actually weak at, so everyone practises the same things and revises what they already know.

The approach. Close the loop: teach, practise, measure, adapt. An AI tutor teaches over chat, voice, and generated video; an exam engine runs real timed attempts with mark-for-review, negative marking, and auto-submit; and what a student gets wrong feeds the next plan. Thirteen role-based career tracks sit on the same engine, with mock interviews, JD analysis, and resume review for students on the way out.

The engineering. Every model call is grounded in structured records, not recall. Profile, active path, and prior tasks load from PocketBase and go in as explicit context; the model is routed to Claude or NVIDIA behind one provider interface; responses are schema-validated before they persist, so off-syllabus output is rejected rather than rendered. Each call is metered against a per-user daily token budget so cost can't run away.

The outcome. 20+ paying users in beta on a $35/month AWS stack. Sub-200KB JS on critical paths, an offline-capable PWA via Serwist, and the same Next.js bundle wrapped as an Android TWA with Bubblewrap — one codebase, Play Store install, works on a weak connection.

  • BackendSelf-hosted PocketBase over hosted Supabase — auth, records, rules, and file storage in one binary on the box already being paid for, with no per-row quota to grow into.
  • Mobile shippingBubblewrap TWA over React Native — same Next.js bundle, no duplicate codebase, Play Store install in under a week.
  • Hosting$35/mo AWS EC2 + nginx + pm2 — predictable cost, no surprise bills, easy to step up to ECS if traffic warrants it.
  • Performance budgetSub-200KB JS on critical paths — measurable, enforceable, falls straight out of Next.js bundle analysis.
  • OfflineSerwist service worker over a native rewrite — the students who need this most are on entry-level Android and patchy data.
20+ paying users in beta$35/mo hosting cost<200KB critical-path JS
Next.js 15React 19TypeScriptTailwindshadcn/uiPocketBase (self-hosted)Anthropic Claude APINVIDIA APIRazorpayResendSentryAWS EC2nginxBubblewrap TWA
Product · Live

HumanifyCV

AI career workspace — a verified career vault that generates, tailors, and ATS-checks resumes

The problem. Resume builders solve templates, which was never the hard part. The hard part is keeping one true record of what someone actually did and re-cutting it per job description without the model quietly inventing an employer, a date, or a metric.

The approach. A structured Career Vault is the source of truth — experience, projects, metrics, education, all typed. The model writes, humanises, and tailors from that vault instead of from its own recall, and a proof-check pass exists to catch claims that aren't backed by it. On top: JD matching, ATS analysis against Workday / Greenhouse / Lever / Taleo / iCIMS parsing heuristics, cover letters, an interview kit, and PDF export. A separate multi-tenant console lets colleges run cohorts, programs, certificates, and missions over the same platform.

The engineering. Auth treated as a feature, not an afterthought: NextAuth v5 with email verification, TOTP 2FA backed by AES-256-GCM-encrypted secrets, WebAuthn passkeys, Google OAuth. Razorpay events modelled as a discriminated union so captures, refunds, and disputes can't silently miscompile. A multi-model Claude router runs the generation.

The outcome. 30–40 paying users on AWS EC2, shipped by GitHub Actions to a Docker Compose stack behind Cloudflare. Sentry for runtime, Amazon SES SMTP for transactional email, Jest / Testing Library tests on the auth + payment paths specifically.

Auth Layerproduction-grade
NextAuth v5Email verificationTOTP 2FAAES-256-GCM secretsWebAuthn passkeys
App LayerAWS EC2
Next.js 16React 19TypeScript strictPrisma 7Postgres
Integration Layertyped boundaries
Claude Sonnet 4.6Razorpay (discriminated union)Nodemailer → SES SMTP
Observability & Qualitynever-silent failures
SentryJest + Testing Library25 test files
  • Payment modelRazorpay events as a TypeScript discriminated union — captured / refunded / disputed can't be confused at the type level.
  • 2FA storageTOTP secrets AES-256-GCM-encrypted at rest with a key from secret manager. Plaintext never touches Postgres.
  • PasskeysWebAuthn FIDO2 over passwords — phishing-resistant, no shared secret, signs in with the device biometric.
  • Test prioritiesTests concentrated on the auth + payment paths — the most damaging failure modes are the regressions caught here first.
30–40 paying users5 ATS engines modelled0 plaintext secrets at rest
Next.js 16React 19TypeScriptPostgres + Prisma 7NextAuth v5WebAuthn passkeysTOTP 2FAAnthropic Claude (multi-model router)RazorpayAWS EC2 + Docker ComposeAmazon SES SMTP (Nodemailer)CloudflareGitHub ActionsSentryJest + Testing Library
Client engagement · Ongoing

AICPA & CIMA Enterprise Platform

Multi-region AWS deployment + label-driven GitOps for a London-based enterprise SaaS client.

The problem. Sequential branch model, manual cherry-picking between release and main, manual SNOW change tickets, ECS deploys done by hand. Release errors were frequent and slow to attribute.

The approach. Replaced the sequential model with a label-driven GitOps topology. Labels on PRs drive the pipeline — adding fbe spawns a full-stack ephemeral environment (ECS + RDS + S3 + SQS + SNS) provisioned via Terraform; adding staging deploys to ECS staging; merging to main auto-opens a ServiceNow Change Request, authenticates to AWS via OIDC keyless, deploys, and closes the CR.

The outcome. ~70% reduction in release errors. PR setup time from days to minutes. Scheduled Friday-night cleanup auto-destroys all FBEs to save weekend spend.

ServiceNow ITSM
CR lifecycle automated end-to-end via REST
Teams ChatOps
Deploy + incident notifications
Trivy + Grype
Shift-left container CVE scanning
  • Trigger modelPR labels — fbe, staging, prod — over branch-per-environment. No drift between branches, label removal cleans up.
  • AWS authOIDC keyless from GitHub Actions to AWS — no long-lived access keys to rotate, no leaked-secret blast radius.
  • Region strategyMulti-region with Route 53 latency routing + active-active failover. Verified failover via game day.
  • Change managementServiceNow CR opened + closed by the pipeline — no manual ticket toil, deploy + CR are atomic.
~70% release errors reduceddays → minutes PR env setup200+ CI/CD pipelines
GitHub ActionsAWS ECSAWS RDSLernaTerraformServiceNowOIDCTrivy / GrypeRoute 53
Bare-metal infra · Production

TimeChamp On-Prem Infrastructure

Production Kubernetes data center built from bare metal at CtrlS Hyderabad. Two years at 99.9% uptime.

The problem. TimeChamp's existing platform was a Windows monolith on IIS, deployed by hand, with no observability and unpredictable cloud spend. Customer growth was capped by release velocity and on-call burden.

The approach. Architected and built a full on-prem Kubernetes data center from bare metal — racked Dell PowerEdge servers at CtrlS Hyderabad, designed VLAN segmentation, configured FortiGate 200F + IPSec VPN with dual-ISP failover. Migrated the monolith to Docker + Kubernetes (~15 nodes, 100+ containers on Hyper-V). Deployed Prometheus + Grafana + Loki for observability, 200+ Azure DevOps pipelines for CI/CD, Cloudflare WAF + CDN at the edge.

The outcome. 99.9% uptime over two years. Release cycle cut ~60% with zero-downtime deployments. $0 cloud spend for on-prem workloads — only burst traffic goes to AWS / Azure.

Physical LayerCtrlS Hyderabad
Dell PowerEdge serversFortiGate 200FIPSec VPN + dual-ISPCisco VLAN segmentation
Orchestration Layer~15 nodes · 100+ containers
KubernetesHyper-VDockerHelm
Observability Layermetrics · logs · alerts
PrometheusGrafanaLoki
CI/CD + Edgeautomation · security
Azure DevOps · 200+ pipelinesCloudflare WAF + CDNDR Validator (C# + AWS S3)
  • On-prem over cloud-nativeDatacenter colo at CtrlS — predictable cost at scale, full network + storage control, regulatory comfort for customer data.
  • Hyper-V under K8sHyper-V for the hypervisor — leverages the team's existing Microsoft expertise and licensing while K8s does the orchestration.
  • Network HAFortiGate 200F + IPSec VPN + dual-ISP failover — verified by pulling the live ISP cable. Zero customer impact.
  • DR ValidatorCustom tool in C# / .NET — restores every MS SQL backup to a throwaway instance daily, alerts on silent corruption.
99.9% uptime · 2 years100+ containers · 15+ nodes$0 cloud spend for on-prem workloads
KubernetesDockerHyper-VFortiGateCiscoPrometheusGrafanaLokiAzure DevOpsCloudflareC# / .NET
FinOps · Cost Engineering

Cloud Cost Engineering

Treating the cloud bill as an SLO. Non-prod that pays rent only when someone's using it, plus the boring commitment work that compounds.

The headline lever. All non-prod AWS scales to zero every Friday at midnight and auto-restores Monday at 6 AM — roughly 54 idle hours a week that used to bill for nothing. It's scheduled automation, so it never depends on someone remembering to shut things down before the weekend.

Monrunning
Tuerunning
Wedrunning
Thurunning
Frirunning
Sat0 replicas
Sun0 replicas
Fri 00:00 → Mon 06:00 · scale-to-zero window
  • Commitment coverageReserved Instances + Savings Plans on the steady-state baseline so predictable load isn't paying on-demand rates.
  • RightsizingInstances matched to real utilization, not the size someone picked once and never revisited.
  • Biweekly reviewsA recurring AWS cost-optimization review — cost is a metric with an owner and a cadence, not a quarterly surprise.
  • Ephemeral cleanupFriday-night teardown of PR environments so "ephemeral" actually means ephemeral.
~54 hrs/wk idle spend eliminated~25% monthly cloud cost cutAICPA & CIMA AWS savings delivered
AWSReserved InstancesSavings PlansScale-to-zero automationRightsizing

Tech Stack

Technologies I use daily to build, ship, and operate production infrastructure and side products.

Cloud (AWS & Azure)

  • AWS (ECS, EKS, RDS, Lambda, S3, Route 53, IAM, multi-region)
  • AWS Bedrock + SageMaker inference endpoints
  • Azure (AKS, Container Apps, VMs / VMSS, Functions, Storage, Azure SQL)
  • Azure Front Door Premium + Application Gateway v2 (WAF_v2)
  • Private Endpoints + Private DNS zones
  • Entra ID, user-assigned managed identity, Key Vault
  • Defender for Cloud + Microsoft Sentinel
  • Azure OpenAI + Azure AI Foundry + Azure Speech
  • On-prem (Hyper-V, bare-metal Kubernetes)
  • GCP (fundamentals)

Containers & Orchestration

  • Kubernetes (production, on-prem + cloud)
  • Docker
  • Helm
  • Kustomize
  • ArgoCD
  • Horizontal Pod Autoscaling
  • Ingress
  • Network policies
  • Admission control
  • cert-manager
  • OpenShift / OKD (self-hosted cluster)

IaC

  • Terraform
  • Bicep (22 authored modules)
  • AWS CDK
  • Ansible

CI/CD & GitOps

  • GitHub Actions (OIDC federated auth)
  • Jenkins
  • Azure DevOps
  • ArgoCD
  • OPA (policy-as-code)

AI Infrastructure

  • Ollama (self-hosted LLMs)
  • DeepSeek / Qwen
  • RAG on pgvector
  • Anthropic Claude API (router: prompt caching, vision)
  • AWS Bedrock (Knowledge Bases, Guardrails, evaluation)
  • Fine-tuned model trained and served on GPU
  • NVIDIA NIM inference microservices
  • LangChain / LlamaIndex
  • Multi-provider LLM router (caching, severity-based cost routing)
  • Prompt engineering
  • AI-in-SDLC auto-remediation

Web & Caching

  • Nginx
  • Redis
  • AWS API Gateway
  • CloudFront

Networking & Security

  • Linux
  • Windows Server (IIS, Hyper-V, MS SQL)
  • Cisco switches
  • FortiGate (firewall, VPN, SD-WAN)
  • IDS / IPS
  • VLAN segmentation
  • Cloudflare WAF
  • Azure WAF tuning (OWASP CRS 3.2, DRS 2.1, anomaly scoring, exemptions)
  • OAuth 2.0 / OIDC with PKCE (Okta)
  • BGP / IPSec
  • OIDC / IAM
  • TLS
  • Trivy / Grype
  • SonarQube
  • ISO 27001:2022 / SOC 2-aligned

Databases

  • PostgreSQL
  • Supabase
  • MS SQL Server
  • MySQL
  • pgvector
  • Azure SQL (CDC, failover groups, dynamic data masking)
  • Debezium + Kafka Connect
  • AWS RDS
  • DynamoDB

Observability & SRE

  • Prometheus
  • Grafana
  • Log Analytics + KQL
  • Application Insights
  • Datadog
  • New Relic
  • CloudWatch
  • Sentry
  • Loki
  • OpenSearch
  • ELK
  • SLOs / DORA metrics
  • Incident response & RCA
  • On-call rotation

Languages & Frameworks

  • TypeScript
  • Python
  • Go
  • C# (.NET)
  • Bash
  • PowerShell
  • SQL / T-SQL
  • KQL
  • React / Next.js
  • Node.js / Express
  • Angular

Engineering Writeups

Longer-form posts on the technical decisions behind my products and infrastructure work. Useful pre-reading for an interview.

● Live

Architecture evolution — three lessons from migrating live systems

IIS→Kubernetes, TeamCity→Azure DevOps, and sequential-branch to label-driven GitOps. What broke, what didn't, and what I'd undo — migrating under live traffic.

anirudhvaka.dev/writeups

○ Planned

Building a production on-prem K8s data center from bare metal

Racking Dell PowerEdge at CtrlS Hyderabad, VLAN segmentation, FortiGate failover, choosing Hyper-V under Kubernetes, and what 99.9% uptime for two years actually cost.

anirudhvaka.dev

Let's build together

Open to AI Infrastructure / MLOps, Platform, and Senior DevOps / SRE roles — relocation worldwide with visa sponsorship, or fully remote. Let's talk.

Serving notice — available from 23 Oct 2026 · open to visa sponsorship / relocation · fully remote-friendly

anirudh.dev shell — type 'help' to list commands.
AvailableSenior DevOps · AI Infrastructure
IST (UTC+5:30)replies < 24hall systems operational