~/himanjan — zsh all systems nominal

$whoami

Himanjan Pati

AI Infrastructure Engineer · LLM performance

Fifteen years keeping production systems up, seven of them as an SRE. Now on the infrastructure side of machine learning — the layer where a model meets real traffic, and where latency, throughput and GPU cost stop being theoretical.

Uptime
15y
Certifications
07
Region
uk-south-1
Inbox
open

whoami --verbose

01

AI infrastructure engineer with fifteen years in production operations, seven of them as a Site Reliability Engineer. I work on the reliability, performance and cost of AI/ML platforms running at scale in a regulated environment — AWS, Kubernetes, and the GPU-backed model serving behind them.

My specialism is the inference layer: what a model costs to serve, where latency originates, and how a serving stack behaves under real concurrency rather than in a demo.

  1. LLM performance benchmarking

    Designed, built and operated a distributed load-testing framework — Locust master and workers on ECS — generating around 10,000 concurrent streaming sessions against models running on A100 and H100 hardware. Measured time to first token, time to first 100 tokens, sustained token throughput and streaming stability under load, and turned the results into capacity and sizing guidance.

    The serving stack was owned by a partner team. The harness, the methodology and the analysis were mine.

  2. Platform reliability

    SLOs, error budgets, incident response and capacity planning for production AI/ML services on EKS, SageMaker, EMR and Databricks — including the on-call rotation that carries them.

  3. Agentic operations tooling

    Platform layer for an agentic incident-triage system: MCP servers, multi-tool orchestration, and the permission model governing what an agent is allowed to do against production systems.

  4. Current focus

    Kubernetes-native inference — vLLM, llm-d and the Gateway API Inference Extension — and the routing, batching and KV-cache decisions that determine cost per million tokens.

I write about this work as I go: the pieces below cover agentic architecture, RAG economics and running MCP servers in production.

Focus
LLM inference infrastructure
Depth
15 yrs production ops · 7 as SRE
Stack
Kubernetes · AWS · GPU serving
Writing
Agentic architecture, RAG economics, MCP in production

cat experience.log

02
2022 — now

JPMorgan Chase & Co.

Bournemouth, UK

AI/ML Site Reliability Engineer

  • Reliability, SLOs and incident response for production AI/ML platforms on AWS — EKS, SageMaker, EMR and Databricks.
  • Designed, built and operated a distributed LLM load-testing framework (Locust master/workers on ECS) driving ~10,000 simulated concurrent users.
  • Benchmarked models served on A100/H100 as the load-testing owner — TTFT, time-to-first-100-tokens, throughput and streaming stability — feeding capacity decisions back to the serving team.
  • Integrated MCP servers and multi-tool orchestration into the platform layer of an agentic incident-triage system.
KubernetesAWSPythonLocustSageMakerTerraformMCP
2021 — 2022

Ageas UK

United Kingdom

DevOps Lead

  • Led the DevOps function for a UK general insurer, owning CI/CD standards and cloud delivery practice.
  • Built reusable Azure Pipelines templates adopted across application teams.
  • Containerised workloads onto AKS and codified infrastructure with Terraform.
  • Embedded security and quality gates directly into the release path.
AzureAKSTerraformAzure DevOpsDocker
2012 — 2021

Infosys Limited

India & UK

Technology Analyst

  • Nine years across the delivery stack — SharePoint SPFx web parts, Azure-hosted web parts, Microsoft 365 and Power Apps development.
  • Designed and shipped enterprise applications for global clients, then moved legacy workloads to the cloud.
  • Grew from developer to technical lead: architecture decisions, code review and mentoring.
SharePoint & SPFxTypeScriptAzureJavaScriptSQL

cat education.log

03
2008 — 2012

B.Tech, Computer Science & Engineering

Biju Patnaik University of Technology · Odisha, India

kubectl get certifications

04
CKA

Certified Kubernetes Administrator

CNCF·2024
CKAD

Certified Kubernetes Application Developer

CNCF·2023
KCNA

Kubernetes & Cloud Native Associate

CNCF·2024
KCSA

Kubernetes & Cloud Native Security Associate

CNCF
SAA

AWS Certified Solutions Architect – Associate

Amazon Web Services·2022
DVA

AWS Certified Developer – Associate

Amazon Web Services·2022
AZ
400

Designing & Implementing Microsoft DevOps Solutions

Microsoft·2021

skills --list --all

05

LLM Inference & Performance

  • Load testing at scale
  • TTFT / TPOT
  • Token throughput
  • Streaming stability
  • vLLM (benchmarked)
  • A100 / H100
  • Continuous batching
  • KV cache
  • Cost per M tokens

Kubernetes & Cloud Native

  • Kubernetes
  • Helm
  • Argo CD
  • Kustomize
  • Docker
  • containerd
  • KEDA
  • Gateway API

Cloud & Data Platforms

  • AWS EKS
  • SageMaker
  • EMR
  • Databricks
  • ECS
  • Lambda
  • Bedrock
  • Azure AKS
  • Azure DevOps

SRE & Observability

  • SLOs & error budgets
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Incident response
  • Capacity planning
  • Postmortems

Agentic Systems

  • MCP
  • Tool calling
  • Multi-tool orchestration
  • RAG architecture
  • Agent evaluation
  • Guardrails
  • LLM gateways & routing

Infrastructure as Code

  • Terraform
  • Ansible
  • GitHub Actions
  • Azure Pipelines
  • Jenkins
  • GitOps

Languages

  • Python
  • Bash
  • TypeScript
  • SQL
  • YAML

ls -lt ~/writing

06

 

 

 

ssh himanjan@contact

07

Happy to talk about reliability for ML systems, Kubernetes at scale, or what LLM workloads really cost once they leave the demo. Drop a note and I’ll come back to you.

Based
Bournemouth, UK · GMT/BST
Reply
usually within a couple of days
Open to
ML reliability, platform engineering, speaking