
- Courses
- Applied AI infrastructure
Applied AI infrastructure
Deploy a production-style AI service on GPU cloud, with an evaluation gate, a cost model, a runbook and a rollback you have demonstrated.

Free
8 modules
6 weeks
Advanced
What you will learn
- Choose between hosted inference and self-hosting from latency, cost and data constraints.
- Size compute, storage and networking, and estimate what the workload will cost.
- Package code, data, model and prompt assets so a build is reproducible.
- Serve at scale with batching, caching, queues, autoscaling and guardrails.
- Gate releases on automated evaluation, and roll back on demand.
- Operate it: traces, drift, GPU utilisation, cost, service objectives and incidents.
Course content
Method step 1, Profile
What the workload actually needs, and what it costs and returns today.
- Training against inference
- Hosted API against self-hosting
- Record latency, throughput and cost today
- Data sensitivity and reliability targets
Method step 2, Build
Sizing the machine, and knowing the bill before it arrives.
- CPU and GPU trade-offs
- Memory and storage patterns
- Data movement, and where it costs
- Cloud, on-premises and hybrid
- Build the cost estimate
Method step 2, Build
A build somebody else can reproduce, including the prompts.
- Environments and containers
- Versioning code, data, model and prompt assets
- Registries and artifact lineage
Method step 2, Build
The serving path, from request to response, under load.
- APIs, batching and caching
- Queues and backpressure
- Kubernetes fundamentals for this workload
- Autoscaling and multi-tenancy
- Secrets and configuration
Method step 2, Build
The choices specific to serving a language or multimodal model.
- Quantization and model selection
- Retrieval and streaming
- Guardrails and fallbacks
Method step 3, Deploy
The service goes to production behind a gate. Deploying it completes the course.
- Continuous integration and delivery
- Evaluation gates
- Release strategies and rollback
- Experiment tracking and change management
- Deploy to production infrastructure
Method step 4, Measure
What you watch, what wakes you, and what the numbers say against module 1.
- Logs, metrics and traces
- Quality evaluation and drift
- GPU utilisation and cost
- Service objectives, alerts and incident response
- Measure against your baseline
Method step 5, Document
The threat model, the evidence, and the record that closes the course.
- Access control and data protection
- Prompt injection and supply chain
- Compliance evidence and vendor inventory
- Demonstrate the rollback
- Assemble the completion record
- Publish it and book the review
Requirements
- One service you own, and permission to change how it is deployed.
- A GPU cloud account, and the ability to spend on it.
- Working knowledge of containers and a CI system.
- About five hours a week for six weeks.
About this course
This course is for developers who own the serving layer, and it is the catalog's advanced anchor. The GMI Cloud work behind it is where these patterns were run at production scale.
It is deliberately compact. The strongest comparable course on the market runs 61 hours, and depth there comes from repeated explanation. Roughly ten to sixteen hours of instruction, with demanding labs and a real evaluation gate, is a more credible claim, so the modules are dense and the labs are where the time goes.
The method is the one every course here runs on. You profile the workload and record what it costs and how it performs today. You build serving, packaging and observability across eight guided labs. You deploy to production infrastructure, then measure latency, utilisation and cost against your baseline.
The capstone is a production-style service with an architecture decision record, a cost model, automated evaluation, a dashboard, a runbook, a threat model, and a rollback you have actually demonstrated.
Common questions
It is free. The first module of every course opens instantly, and one free account unlocks the rest of that course. Everything on this site is free to use.
A course completes when your workflow runs live and you have measured it. You record a baseline in module 1, launch the workflow into a working environment, then measure the same thing again. That pairing is what makes the completion record worth sharing.
Yes. The first module of every course plays for everyone, with a free account needed only from module 2.
Name, role, organization, and a few questions about the workflow you want to improve. It takes about a minute, and it is what lets the guided labs use your own workflow as the project.
One workflow you own, permission to change it, and the tools your team already uses. Applied AI infrastructure also assumes access to a GPU cloud account, and AI starter for small business runs on everyday business tools.
A one-page record inside the course. Before you launch, you note a baseline on time, cost, or quality. After you launch, you measure again. The sheet holds both numbers and the difference between them.
Instructor

Roan Weigert
Records all five courses
Hackathon judge and AI content creator, and host of the AI Insights San Francisco podcast. Writes this curriculum and records every core lesson.
Other courses
The other four courses
Each one takes a different kind of work through the same five steps. Pick the one closest to your job.
You buildAn agent workflow on your live pipeline
You buildAn ingest and rough-cut pipeline
You buildAn AI use policy your team runs onAI literacy, ethics and data compliance
A working AI policy for your team, covering what the tools may touch and who checks it.
You buildOne assistant for your busiest taskAI starter for small business
A short course to one measurable first win, built for owners and operators.
Start with module 1
Architecture and workload discovery
Open to everyone, with no account. You finish it holding a baseline and decision record.
