Production AI Infrastructure
A Meta-approved SaaS platform handling real client traffic since April 2026. Not a tutorial project — live in production.
See it in productionCloud Infrastructure & DevOps Engineer
Building and operating production AI systems on AWS, Kubernetes, and Terraform with a security-first, least-privilege approach at every layer.
I design, deploy, and operate production cloud infrastructure with a security-first, least-privilege approach at every layer.
A Meta-approved SaaS platform handling real client traffic since April 2026. Not a tutorial project — live in production.
See it in productionInfrastructure, networking, application layer, DevOps pipeline, and security scanning. From the first line of Terraform to production traffic.
See it in production15 years running multi-location businesses before engineering. I understand what downtime actually costs the bottom line.
See it in productionAverage DM response time. Down from 6 hours on the same workload.
DM-to-booking conversion rate on live production traffic.
CI/CD build time. Reduced from 20 minutes with caching and pipeline optimization.
Terraform-managed AWS resources deployed during a live migration with zero downtime.
Real systems handling real traffic. Every diagram below maps to infrastructure running today.
Featured production project
Meta-approved multi-tenant AI SaaS automating Instagram DM responses for beauty and wellness businesses. Built the complete stack from scratch — infrastructure, networking, application layer, and DevOps pipeline. The AI receptionist responds in under 60 seconds, handles booking intent, and hands off to an agentic Playwright booking agent for live availability lookups. First client: Secretive Nail Bar, three Southern California locations.
Key decisions: Provider abstraction layer bridges Claude API and AWS Bedrock for runtime switching without code changes. Agentic Playwright agent runs on a separate ECS task triggered by booking intent. Split CI/CD pipelines with path-based triggers so app changes never touch infrastructure pipelines.
Platform engineering project
Kubernetes-based prospect enrichment and intelligence platform spanning AWS and Azure. A Playwright agent aggregates and enriches real business data via Google Places API, scored and surfaced through a dashboard. Built on a hub-and-spoke multi-cloud topology with GitOps via ArgoCD, canary deployments via Argo Rollouts, and Vault sidecar secret injection.
Pipeline logic: Every git push triggers GitHub Actions for build and security scanning. Helm packages the manifests. ArgoCD detects the new image tag and syncs. Argo Rollouts splits traffic at 20% canary. Health checks gate promotion. Failed probes trigger automatic rollback with no manual intervention.
Pod and instance level metrics. Custom dashboards per workload.
Platform agnostic by design. Azure ACA planned for next phase.
Secrets never in manifests or Git. Rotated independently of deployments.
Design principle: Infrastructure and application layers are fully decoupled. Terraform manages the platform. ArgoCD manages the workloads. Vault manages secrets. Each layer is independently replaceable without touching the others.
Built and operate a Meta-approved production AI SaaS platform on AWS, live client traffic, real business outcomes, zero production outages since launch.
Reduced average DM response time from 6 hours to under 60 seconds via event-driven architecture: API Gateway, SQS decoupling, ECS Fargate consumer.
Cut CI/CD pipeline build time from 20 minutes to 5 minutes using Docker layer caching, path-based triggers, and split app and infra pipelines.
Implemented GitOps with ArgoCD and Argo Rollouts, canary deployments, automated rollback on health check failure, Git as the single source of truth.
Hardened CI/CD security pipeline with Gitleaks secrets scanning, Bandit SAST, Trivy image scanning, and pip-audit required before every merge.
Designed and deployed an agentic AI booking integration using Playwright on managed ECS compute, triggered by booking intent detection in live conversation flow.
Systematic production debugging approach: establish blast radius first, isolate to network, application, or infrastructure layer, trace through logs and exit codes, remediate, then document root cause and update runbooks to prevent recurrence.
Designed and deployed a full observability stack using Prometheus sidecar injection and Grafana dashboards with custom pod and instance level metrics, moving beyond default cluster metrics to measure what actually matters for the application workload.
I am a Cloud Infrastructure and DevOps Engineer focused on AWS, Kubernetes, Terraform, and production AI systems. I design and operate infrastructure end-to-end with a security-first approach at every layer.
Before engineering, I spent 15 years running multi-location businesses. That background changed how I think about infrastructure. I know what it feels like from the stakeholder side when systems fail to meet objectives. I am not task-oriented. I am outcome-oriented at my core.
I am comfortable under pressure, meticulous about security, and genuinely enjoy debugging complex flows when things break. I combine operational discipline with clear, direct communication across technical and non-technical teams.
I have always been drawn to solving complex problems. Cloud and DevOps gave me the perfect outlet to realize that characteristic.
Open to Cloud Engineer, DevOps Engineer, Platform Engineer, and SRE roles. Remote, hybrid, or on-site. Costa Mesa, CA. Contract, contract-to-hire, or full time.