Skip to content
AI Tech

Applied AI infrastructure

Deploy a production-style AI service on GPU cloud, with an evaluation gate, a cost model, a runbook and a rollback you have demonstrated.

Created by Roan Weigert

Updated August 2026·English·Self-paced

Free

Every module

  • 8 modules

    Lessons and labs throughout

  • 6 weeks

    At your own pace

  • Advanced

    Recommended

What you will learn

  • Choose between hosted inference and self-hosting from latency, cost and data constraints.
  • Size compute, storage and networking, and estimate what the workload will cost.
  • Package code, data, model and prompt assets so a build is reproducible.
  • Serve at scale with batching, caching, queues, autoscaling and guardrails.
  • Gate releases on automated evaluation, and roll back on demand.
  • Operate it: traces, drift, GPU utilisation, cost, service objectives and incidents.
  • Model serving
  • GPU orchestration
  • Cost control
  • Observability

Course content

8 modules · 36 lessons · guided labs throughout

Method step 1, Profile

What the workload actually needs, and what it costs and returns today.

  • Training against inference
  • Hosted API against self-hosting
  • Record latency, throughput and cost today
  • Data sensitivity and reliability targets

You finish with Baseline and decision record

Requirements

  • One service you own, and permission to change how it is deployed.
  • A GPU cloud account, and the ability to spend on it.
  • Working knowledge of containers and a CI system.
  • About five hours a week for six weeks.

About this course

This course is for developers who own the serving layer, and it is the catalog's advanced anchor. The GMI Cloud work behind it is where these patterns were run at production scale.

It is deliberately compact. The strongest comparable course on the market runs 61 hours, and depth there comes from repeated explanation. Roughly ten to sixteen hours of instruction, with demanding labs and a real evaluation gate, is a more credible claim, so the modules are dense and the labs are where the time goes.

The method is the one every course here runs on. You profile the workload and record what it costs and how it performs today. You build serving, packaging and observability across eight guided labs. You deploy to production infrastructure, then measure latency, utilisation and cost against your baseline.

The capstone is a production-style service with an architecture decision record, a cost model, automated evaluation, a dashboard, a runbook, a threat model, and a rollback you have actually demonstrated.

Common questions

Instructor

Portrait of Roan Weigert

Roan Weigert

Developer Relations AI Engineer, San Francisco

Records all five courses

Hackathon judge and AI content creator, and host of the AI Insights San Francisco podcast. Writes this curriculum and records every core lesson.

Other courses

The other four courses

Each one takes a different kind of work through the same five steps. Pick the one closest to your job.

Compare every course

Start with module 1

Architecture and workload discovery

Open to everyone, with no account. You finish it holding a baseline and decision record.

Enroll for freeStarts Aug 10

A free account opens the rest.