LLM APIs and Prompting · Calling a Model · lesson 3 of 8
System prompts and instruction design
about 15 minutes · free · runs in your browser
Step 1 of 2
Standing instructions
The system prompt is where the rules live: the role, the format, the things never to do. It is sent on every call, which is what makes it standing rather than a one-off request buried in a user message.
A system prompt earns its length by being specific about output. "Be helpful" changes nothing measurable. "Answer in one sentence, with no preamble" changes the shape of every reply, and the shape is what your code has to parse.
fake_llm.Client holds the defaults so they are not retyped at every call site — the
same reason you would wrap a real SDK client rather than passing the same three arguments
everywhere.
client = fake_llm.Client(system="Answer in one sentence.", temperature=0)
client.ask("What is an embedding?") # the text, straight back
Your turn: build make_client(rule) returning a client whose system prompt is
rule, and check with the call log that the rule really is being sent — on every call,
not just the first.
You start from this, and edit it in the browser:
import fake_llm
def make_client(rule):
"""Return a Client that sends 'rule' as its system prompt."""
return None
def system_sent(client, question):
"""Ask the question, then report the system message the API received."""
return ""
Step 2 of 2
Asking for a shape, not a vibe
The useful instructions are the ones with an observable consequence. Compare:
- "Classify the sentiment of the review." — you get a paragraph about the review.
- "Reply with exactly one word: positive, negative or neutral." — you get a word your code can branch on.
The second is not politer. It is testable, which is the whole difference between a prompt you can ship and a prompt you can only admire. And it fails loudly: if the model answers with a sentence, your parse raises, and you know immediately rather than three screens later when something downstream gets a paragraph where it expected a label.
Your turn: write classify(review) returning "positive", "negative" or
"neutral". Instruct the model with a system prompt that asks for the sentiment as a
single word, then normalise what comes back — strip the whitespace and lowercase it,
because a model that was told "one word" will still occasionally shout.
You start from this, and edit it in the browser:
import fake_llm
RULE = "Say something about the review."
def classify(review):
"""Return 'positive', 'negative' or 'neutral' for one review."""
return ""