Usama Zaheer
I design, automate and run infrastructure across AWS, Azure and GCP: Kubernetes platforms, Terraform, GitOps delivery, GPU inference fleets and the cost controls that keep them affordable. Seven years in, based in Lahore.
| NAME | READY | STATUS | AGE |
|---|---|---|---|
| stealth-startup | 1/1 | Running | 1m |
| alethia-ai | 1/1 | Running | 2y10m |
| data-pilot | 1/1 | Completed | 2y2m |
| venture-data | 1/1 | Completed | 2y4m |
usama@infra:~/experience$ cat *.logExperience
DevOps Lead, Senior DevOps / MLOps Engineer
Owns cloud, Kubernetes and GPU infrastructure end to end and mentors a direct report.
Platform and delivery
- Ran Kubernetes on EKS, AKS, GKE and K3s with namespace isolation and resource quotas for multi-tenant workloads.
- GitOps delivery with Argo CD and Helm; CI/CD on GitHub Actions, Azure DevOps Pipelines and Jenkins.
- Migrated a legacy backend from EC2 and CodeDeploy to EKS with the whole stack in Terraform, OIDC-based CI with no static keys, and per-pod IAM roles.
- Implemented AWS IAM Identity Center as policy as code across an 11-account AWS Organization.
GPU and MLOps
- Kept real-time streaming inference on RTX 4090s running across Vast.ai, RunPod and a local fallback through a market-wide GPU shortage.
- Designed GPU autoscaling with time-slicing, priority classes, namespace quotas and Prometheus-driven KEDA scaling.
- Shortlisted vendors across 10+ GPU clouds and secured credits from RunPod, SaladCloud and BytePlus.
- Worked with Triton Inference Server, TorchServe and MLflow.
FinOps and security
- Reduced AWS spend from $25,000 to $7,000 a month through reserved capacity, Spot, right-sizing and idle-resource cleanup.
- Built cost tracking for 89 SaaS and AI subscriptions with a named owner per line, plus CEO-facing AI spend reports.
- Shipped per-product OpenAI and Anthropic spend alerts to Slack as a Kubernetes CronJob.
- Hardened email with SPF
-all, DKIM and DMARCp=rejectacross Google Workspace and Amazon SES.
DevOps Engineer
- Designed and ran Azure and AWS infrastructure for data engineering and ML teams across enterprise clients.
- Delivered on the Azure DevOps suite (Repos, Pipelines, Boards, Artifacts), Azure API Management and Microsoft Fabric.
- Built and maintained MLflow infrastructure for Nokia Saudi Arabia.
- Ran analytics infrastructure for Lulusar (Metabase on AWS), Victoria Beckham Beauty, Tripletree and Interloop.
- Deployed ETL on AWS Lambda and Glue into Redshift, and orchestrated workflows with Airflow.
- Deployed an on-premises people-tracking computer vision system for a department store and ran an Azure-hosted food inspection CV system for Kitchen Cuisine.
- Set up observability with Prometheus, Grafana and custom Python tooling.
DevOps Engineer
- Managed AWS and GCP environments for multiple clients, including migrations and disaster recovery.
- Saved $7,000 a month, about 30% of infrastructure cost, through EC2 right-sizing, storage changes and database tuning.
- Configured WAF ACLs, IAM policies, federated access and VPNs.
- Built QuickSight dashboards on Glue, Athena and Kinesis for near real-time analysis.
- Centralized logs and alerting with CloudWatch, CloudTrail and Datadog.
usama@infra:~/work$ git log --stat changes/Selected work
Kept live GPU streams up through an RTX 4090 shortage
- problem:
- Two real-time lip-sync and voice services ran on rented 4090s. Vast.ai prices spiked and capacity disappeared.
- change:
- Moved both services to RunPod with a local 4090 as a warm fallback, and priced a RunPod savings plan for leadership.
- result:
- Services stayed up, with credits secured from RunPod, SaladCloud and BytePlus.
Moved a legacy backend from EC2 to EKS, all in Terraform
- problem:
- An older product ran on EC2 with CodeDeploy and static CI credentials.
- change:
- Terraform for the cluster, node groups, IAM, OIDC, EFS and EBS CSI, Load Balancer Controller and IRSA. CI split into 6 jobs with migrations gated as a stage.
- result:
- No long-lived keys in CI, least-privilege roles per pod, about 800 GB of existing media reused without a cutover, and a $196 a month cost baseline.
Put every AI and SaaS bill under a named owner
- problem:
- Spend was shifting from AWS to AI APIs and tool seats with no clear ownership.
- change:
- A tracker for 89 subscriptions with owners and cost history, monthly variable cost reports for the CEO, and a dedicated API key per product with alerts at 50, 80 and 100% of budget.
- result:
- Spend attributable per product, with hourly checks posting to Slack.
Per-person AWS access as code
- problem:
- A startup where each person needs different access, across an 11-account AWS Organization that keeps growing.
- change:
- IAM Identity Center permission sets and assignments defined in Terraform at the individual level, not only by team.
- result:
- Access changes go through review, and new accounts slot into the same model.
Traced 10 of 13 crash loops to one metadata setting
- problem:
- New Spot node groups on staging left most pods in CrashLoopBackOff.
- change:
- Found the groups had no launch template, so the IMDS hop limit defaulted to 1 while pods sit two hops out. Rebuilt both groups with a launch template setting it to 2.
- result:
- Staging back on cheaper Spot capacity, with the two unrelated failures isolated for their owners.
A sandboxed worker tier for untrusted jobs
- problem:
- A media pipeline ran AI agent workers next to everything else on a shared cluster.
- change:
- A dedicated node pool provisioned by Karpenter, gVisor for sandboxing, and KEDA scaling the workers from an SQS queue.
- result:
- Agent workloads separated from the rest of the cluster and scaled on queue depth.
usama@infra:~$ cat stack.yamlStack
usama@infra:~$ ls -l certs/Credentials
| -r--r--r-- | Amazon Web Services | aws-solutions-architect-associate | |
| -r--r--r-- | Microsoft | azure-fundamentals | |
| -r--r--r-- | LinkedIn Learning | Jun 2023 | learning-kubernetes |
| -r--r--r-- | LinkedIn Learning | Jun 2023 | learning-terraform |
| -r--r--r-- | Islamia Univ. Bahawalpur | 2017 – 2021 | bba-finance |
usama@infra:~/recommendations$ cat *.txtRecommendations
# 5 recommendations, from LinkedIn
Usama is one of the most capable DevOps and Cloud Architects I have worked with. From architecting scalable multi-cloud Kubernetes clusters to designing end-to-end MLOps pipelines and cutting cloud spend by 20%+, his impact is immediate and measurable.
He combines deep technical mastery of AWS, GCP, and IaC with a proactive, reliable work ethic. If you are building high-performance, secure, and cost-effective cloud or AI infrastructure, Usama is the engineer you want on your team.
I had the pleasure of working alongside Usama Zaheer at ALETHIA AI, and he is hands-down one of the most dedicated and skilled DevOps engineers I’ve collaborated with.
Usama has a deep mastery of the entire DevOps and MLOps ecosystem. Whenever high-priority production issues arose, he was always the first to jump in available virtually 24/7 to step up, troubleshoot, and solve complex system bottlenecks quickly.
I had the pleasure of working with Usama as a Senior DevOps Engineer, and I can confidently recommend him as a highly capable and dependable professional. He is a quick learner, an efficient problem solver, and someone who takes ownership of challenges.
His work on optimizing infrastructure and reducing costs brought significant value to the organization. Usama is also willing to take calculated risks when needed and consistently focuses on delivering practical, effective solutions.
It was an absolute pleasure to work with Usama Zaheer. During our time working together at Data Pilot, Usama proved himself to be a highly skilled and incredibly hardworking DevOps Engineer.
Usama has a remarkable talent for optimizing infrastructure and streamlining deployment processes. His ability to build robust, adaptable CI/CD pipelines and his proactive approach to problem-solving significantly enhanced the reliability and efficiency of our systems.
I had the pleasure of working alongside Usama at Data Pilot, and I can confidently say they are one of the most skilled and dedicated DevOps engineers I've encountered. His ability to optimize our infrastructure and streamline deployment processes was instrumental in enhancing the reliability and efficiency of our systems.
Usama has a remarkable knack for automation and problem-solving, ensuring that our CI/CD pipelines were robust and adaptable to changes. I was continually impressed by his proactive approach to identifying potential issues before they could impact production.
usama@infra:~$ cat .contactContact
# Open to conversations about platform engineering, Kubernetes, GPU infrastructure and cloud cost reviews.