Call a model from Python, and pay attention to what it costs and how it fails.
8 lessons · about 6 hours · free
An LLM is an HTTP API that happens to answer in prose. This course covers the call itself — the messages list, system prompts, temperature and the other sampling controls — and then the two things that separate a demo from a program people rely on: knowing what each call costs in tokens and money, and handling rate limits, timeouts and truncated answers without losing the request. By the end you will have written a cost meter, a streaming reader and a retry loop with exponential backoff.
The request, the response, and the four settings that change what comes back.
What every call costs, how to read an answer as it arrives, and what to do when the API says no.