"infra"

Usama Zaheer

usama@stealth-startup
RoleSenior Forward Deployed Engineer - DevOps & Infrastructure
HostStealth Startup (United States, remote)
LocationLahore, Pakistan
Uptime7 years
CloudsAWS, Azure, GCP
KubernetesEKS, AKS, GKE, K3s
IaCTerraform, CloudFormation
DeliveryArgo CD, Helm, GitHub Actions, Azure DevOps
GPU / MLTriton, MLflow, RunPod, Vast.ai
Shellzsh

I design, automate and run infrastructure across AWS, Azure and GCP: Kubernetes platforms, Terraform, GitOps delivery, GPU inference fleets and the cost controls that keep them affordable. Seven years in, based in Lahore.

aws-spend
-72%
$25,000 to $7,000 a month
release-cycle
-50%
after CI/CD rework
venture-data
$7,000
saved per month by right-sizing
gpu-vendors
10+
GPU clouds evaluated and shortlisted
NAMEREADYSTATUSAGE
stealth-startup1/1Running1m
alethia-ai1/1Running2y10m
data-pilot1/1Completed2y2m
venture-data1/1Completed2y4m

usama@infra:~/experience$ cat *.logExperience

[ stealth-startup ]● running

Senior Forward Deployed Engineer - DevOps & Infrastructure

Stealth Startup · Full-timeOct 2026 – PresentUnited States, remote
[ alethia-ai ]● running

DevOps Lead, Senior DevOps / MLOps Engineer

Alethia AI · ContractJan 2024 – PresentLahore

Owns cloud, Kubernetes and GPU infrastructure end to end and mentors a direct report.

Platform and delivery

  • Ran Kubernetes on EKS, AKS, GKE and K3s with namespace isolation and resource quotas for multi-tenant workloads.
  • GitOps delivery with Argo CD and Helm; CI/CD on GitHub Actions, Azure DevOps Pipelines and Jenkins.
  • Migrated a legacy backend from EC2 and CodeDeploy to EKS with the whole stack in Terraform, OIDC-based CI with no static keys, and per-pod IAM roles.
  • Implemented AWS IAM Identity Center as policy as code across an 11-account AWS Organization.

GPU and MLOps

  • Kept real-time streaming inference on RTX 4090s running across Vast.ai, RunPod and a local fallback through a market-wide GPU shortage.
  • Designed GPU autoscaling with time-slicing, priority classes, namespace quotas and Prometheus-driven KEDA scaling.
  • Shortlisted vendors across 10+ GPU clouds and secured credits from RunPod, SaladCloud and BytePlus.
  • Worked with Triton Inference Server, TorchServe and MLflow.

FinOps and security

  • Reduced AWS spend from $25,000 to $7,000 a month through reserved capacity, Spot, right-sizing and idle-resource cleanup.
  • Built cost tracking for 89 SaaS and AI subscriptions with a named owner per line, plus CEO-facing AI spend reports.
  • Shipped per-product OpenAI and Anthropic spend alerts to Slack as a Kubernetes CronJob.
  • Hardened email with SPF -all, DKIM and DMARC p=reject across Google Workspace and Amazon SES.
[ data-pilot ]completed

DevOps Engineer

Data Pilot, data and AI consultancyDec 2021 – Jan 2024Lahore, hybrid
  • Designed and ran Azure and AWS infrastructure for data engineering and ML teams across enterprise clients.
  • Delivered on the Azure DevOps suite (Repos, Pipelines, Boards, Artifacts), Azure API Management and Microsoft Fabric.
  • Built and maintained MLflow infrastructure for Nokia Saudi Arabia.
  • Ran analytics infrastructure for Lulusar (Metabase on AWS), Victoria Beckham Beauty, Tripletree and Interloop.
  • Deployed ETL on AWS Lambda and Glue into Redshift, and orchestrated workflows with Airflow.
  • Deployed an on-premises people-tracking computer vision system for a department store and ran an Azure-hosted food inspection CV system for Kitchen Cuisine.
  • Set up observability with Prometheus, Grafana and custom Python tooling.
[ venture-data ]completed

DevOps Engineer

Venture DataAug 2019 – Nov 2021Lahore, hybrid
  • Managed AWS and GCP environments for multiple clients, including migrations and disaster recovery.
  • Saved $7,000 a month, about 30% of infrastructure cost, through EC2 right-sizing, storage changes and database tuning.
  • Configured WAF ACLs, IAM policies, federated access and VPNs.
  • Built QuickSight dashboards on Glue, Athena and Kinesis for near real-time analysis.
  • Centralized logs and alerting with CloudWatch, CloudTrail and Datadog.

usama@infra:~/work$ git log --stat changes/Selected work

commit CHG-01gpu · alethia-ai

Kept live GPU streams up through an RTX 4090 shortage

problem:
Two real-time lip-sync and voice services ran on rented 4090s. Vast.ai prices spiked and capacity disappeared.
change:
Moved both services to RunPod with a local 4090 as a warm fallback, and priced a RunPod savings plan for leadership.
result:
Services stayed up, with credits secured from RunPod, SaladCloud and BytePlus.
vast.airunpodsaladcloudcuda
commit CHG-02migration · alethia-ai

Moved a legacy backend from EC2 to EKS, all in Terraform

problem:
An older product ran on EC2 with CodeDeploy and static CI credentials.
change:
Terraform for the cluster, node groups, IAM, OIDC, EFS and EBS CSI, Load Balancer Controller and IRSA. CI split into 6 jobs with migrations gated as a stage.
result:
No long-lived keys in CI, least-privilege roles per pod, about 800 GB of existing media reused without a cutover, and a $196 a month cost baseline.
eksterraformirsagithub-oidc
commit CHG-03finops · alethia-ai

Put every AI and SaaS bill under a named owner

problem:
Spend was shifting from AWS to AI APIs and tool seats with no clear ownership.
change:
A tracker for 89 subscriptions with owners and cost history, monthly variable cost reports for the CEO, and a dedicated API key per product with alerts at 50, 80 and 100% of budget.
result:
Spend attributable per product, with hourly checks posting to Slack.
k8s-cronjobecrslackfastapidynamodb
commit CHG-04access · alethia-ai

Per-person AWS access as code

problem:
A startup where each person needs different access, across an 11-account AWS Organization that keeps growing.
change:
IAM Identity Center permission sets and assignments defined in Terraform at the individual level, not only by team.
result:
Access changes go through review, and new accounts slot into the same model.
iam-identity-centerterraformaws-organizations
commit CHG-05incident · alethia-ai

Traced 10 of 13 crash loops to one metadata setting

problem:
New Spot node groups on staging left most pods in CrashLoopBackOff.
change:
Found the groups had no launch template, so the IMDS hop limit defaulted to 1 while pods sit two hops out. Rebuilt both groups with a launch template setting it to 2.
result:
Staging back on cheaper Spot capacity, with the two unrelated failures isolated for their owners.
eksspot-r6aimdsv2
commit CHG-06isolation · alethia-ai

A sandboxed worker tier for untrusted jobs

problem:
A media pipeline ran AI agent workers next to everything else on a shared cluster.
change:
A dedicated node pool provisioned by Karpenter, gVisor for sandboxing, and KEDA scaling the workers from an SQS queue.
result:
Agent workloads separated from the rest of the cluster and scaled on queue depth.
karpentergvisorkedasqs
drwxr-xr-xrag-chatbot/A starter for chatting over your own documents: FastAPI, Postgres with pgvector, and a Next.js front end.
drwxr-xr-xdfir-lab/A Timesketch forensics stack on Docker Compose with OpenSearch, Postgres, Redis and a custom ingestor, run in a home lab VM.

usama@infra:~$ cat stack.yamlStack

cloud:AWS, Azure, GCP, DePIN compute
kubernetes:EKS, AKS, GKE, K3s, Helm, Karpenter, KEDA, gVisor
iac:Terraform, CloudFormation
ci_cd:Argo CD, Argo Workflows, GitHub Actions, Azure DevOps, Jenkins
observability:Prometheus, Grafana, Loki, CloudWatch, Datadog
ml_gpu:Triton, TorchServe, MLflow, CUDA, RunPod, Vast.ai, SaladCloud, CoreWeave
data:Airflow, Glue, Redshift, Athena, Kinesis, Microsoft Fabric, Metabase
security:IAM Identity Center, IRSA, OIDC, SCPs, WAF, SPF / DKIM / DMARC

usama@infra:~$ ls -l certs/Credentials

-r--r--r--Amazon Web Servicesaws-solutions-architect-associate
-r--r--r--Microsoftazure-fundamentals
-r--r--r--LinkedIn LearningJun 2023learning-kubernetes
-r--r--r--LinkedIn LearningJun 2023learning-terraform
-r--r--r--Islamia Univ. Bahawalpur2017 – 2021bba-finance

usama@infra:~/recommendations$ cat *.txtRecommendations

# 5 recommendations, from LinkedIn

[ aamir-naveed ]2026-09-09

Usama is one of the most capable DevOps and Cloud Architects I have worked with. From architecting scalable multi-cloud Kubernetes clusters to designing end-to-end MLOps pipelines and cutting cloud spend by 20%+, his impact is immediate and measurable.

He combines deep technical mastery of AWS, GCP, and IaC with a proactive, reliable work ethic. If you are building high-performance, secure, and cost-effective cloud or AI infrastructure, Usama is the engineer you want on your team.

Aamir NaveedPlatform Engineer, DevOps, SREmanaged Usama directly
[ ali-wisam ]2026-09-12

I had the pleasure of working alongside Usama Zaheer at ALETHIA AI, and he is hands-down one of the most dedicated and skilled DevOps engineers I’ve collaborated with.

Usama has a deep mastery of the entire DevOps and MLOps ecosystem. Whenever high-priority production issues arose, he was always the first to jump in available virtually 24/7 to step up, troubleshoot, and solve complex system bottlenecks quickly.

Ali WisamPrincipal Blockchain Engineer & Quant Developerworked together at Alethia AI
[ zain-shahid ]2026-09-15

I had the pleasure of working with Usama as a Senior DevOps Engineer, and I can confidently recommend him as a highly capable and dependable professional. He is a quick learner, an efficient problem solver, and someone who takes ownership of challenges.

His work on optimizing infrastructure and reducing costs brought significant value to the organization. Usama is also willing to take calculated risks when needed and consistently focuses on delivering practical, effective solutions.

Zain ShahidSenior Software Engineerworked on the same team
[ m-usama-irfan ]2026-06-09

It was an absolute pleasure to work with Usama Zaheer. During our time working together at Data Pilot, Usama proved himself to be a highly skilled and incredibly hardworking DevOps Engineer.

Usama has a remarkable talent for optimizing infrastructure and streamlining deployment processes. His ability to build robust, adaptable CI/CD pipelines and his proactive approach to problem-solving significantly enhanced the reliability and efficiency of our systems.

Muhammad Usama IrfanHelping enterprises build scalable AI solutionsworked together at Data Pilot
[ m-irfan-umar ]2024-11-08

I had the pleasure of working alongside Usama at Data Pilot, and I can confidently say they are one of the most skilled and dedicated DevOps engineers I've encountered. His ability to optimize our infrastructure and streamline deployment processes was instrumental in enhancing the reliability and efficiency of our systems.

Usama has a remarkable knack for automation and problem-solving, ensuring that our CI/CD pipelines were robust and adaptable to changes. I was continually impressed by his proactive approach to identifying potential issues before they could impact production.

Muhammad Irfan UmarBuilding Scalable Data Solutionsworked together at Data Pilot

usama@infra:~$ cat .contactContact

# Open to conversations about platform engineering, Kubernetes, GPU infrastructure and cloud cost reviews.

email=usamazaheer222@gmail.com