Product & Model Selection
There is no single "use Claude" button. There is a surface to work on (chat, a Project, research mode, an artifact) and a model to do the work: Haiku, Sonnet, or Opus. Most people pick both by habit and never notice the cost. This lab teaches the two decisions that separate someone who uses Claude from someone who deploys it well: matching the surface to the shape of the work, and matching the model to the shape of the task, then keeping a conversation inside the limits it can actually reason across.
- Select appropriate Claude product features (Projects, research mode, chat, artifacts)
- Differentiate between Claude model types (Haiku, Sonnet, Opus)
- Align model selection with task requirements (cost, speed, quality)
- Understand and manage context limitations and memory considerations (when to restart, summarize, or persist)
What you need: a Claude account (claude.ai) and about 40 minutes. Everything here happens inside the Claude product: no installation, no API key, no code.
1. Choosing the Product Surface
Before you type a word, you have already made a choice: where the work lives. The same request behaves differently in a throwaway chat than in a Project with standing instructions, and picking the wrong surface is a quiet tax you pay on every turn afterward. The skill is not knowing what each feature is; it is recognizing, from the shape of the work in front of you, which one the situation calls for.
| The situation | Use | Why this one |
|---|---|---|
| A quick, one-off question or a bit of exploration you won't return to | Plain chat | Zero setup, zero overhead. There is nothing to persist and nothing to reuse, so structure would only slow you down. |
| Recurring work that keeps needing the same context, standards, or reference material | Projects | Persistent instructions and uploaded knowledge apply to every conversation inside the Project, so you stop re-pasting the same background and get consistent output across sessions. |
| A question that spans many sources and needs them gathered, weighed, and synthesized with citations | Research mode | It investigates across multiple sources and returns a synthesis you can trace back to where each claim came from, work that a single chat answer cannot do credibly. |
| A substantial deliverable you will iterate on, revise, and reuse — a document, a plan, a piece of writing | Artifacts | The deliverable becomes a standalone object you can refine turn after turn without it scrolling away into the conversation, and you can carry it forward. |
Read the table as a decision, not a menu. The tell is almost always persistence and reuse: if you will need this again (the same instructions, the same knowledge, the same evolving document), a plain chat is the wrong surface, because a chat is something you will lose. If you genuinely won't, then a Project is ceremony you don't need.
- Take a task you did last week in a plain chat that you have since redone or referred back to: a recurring report, a standard email, a document you keep revising.
- Ask yourself the diagnostic question: did I re-supply the same background the second time? If yes, that work wanted a Project.
- Create a Project, move the standing context into its instructions, and run the task again. Notice how much shorter your actual prompt becomes when the background is already there.
- Now take a genuinely one-off question and answer it in plain chat. Feel the difference: the right surface is the one that adds exactly as much structure as the work rewards, and no more.
2. The Three Model Families
Underneath the surface, you are choosing a model. Claude comes in three families, and the exam expects you to reason about them as a single trade-off with three dials: cost, speed, and quality. You cannot maximize all three at once. Every choice buys one by spending another.
| Task profile | Model | Why it fits |
|---|---|---|
| Straightforward, well-defined, high-volume work where speed and cost matter more than deep reasoning — classification, short replies, quick extraction, simple formatting at scale | Haiku | Fastest and lowest cost. The task doesn't require heavy reasoning, so paying for it would be waste; volume makes speed and cost the dominant concern. |
| Most everyday professional work — drafting, analysis, summarizing, editing, general problem-solving where you want strong quality without paying top rates | Sonnet | The balanced default. It carries the bulk of real work with strong quality at moderate cost and speed. When you're unsure, start here. |
| Genuinely complex reasoning — multi-step analysis, subtle judgment, ambiguous problems, high-stakes output where a mistake is expensive | Opus | Most capable. It handles depth and nuance the others may miss, but it costs more and runs slower, so you spend that budget only where the difficulty actually demands it. |
Notice the columns move together. As you go from Haiku to Opus you climb the quality dial and spend more cost and more latency to do it. The question is never "which model is best" in the abstract; it is "which point on this trade-off does this task justify."
3. The Core Insight: The Best Model Is Not the Safe Default
Here is the mistake the exam is built to catch, and it is worth stating plainly: reaching for the most capable model on every task is not caution, it is waste. Opus on a task that Haiku would nail spends cost and latency budget to buy quality the task never needed. Across a handful of chats you won't notice. Across a team, a workflow, or a high-volume queue, that habit is expensive and slow for no return.
The opposite error is just as real. Point the cheapest, fastest model at nuanced reasoning and you get output that looks fine and is subtly wrong, which you then have to catch, discard, and redo. The cheap answer you have to redo was never cheap.
Work the blueprint's own example. You need to generate a high volume of short customer-reply drafts. The replies are short and fairly routine; what matters is turning them around fast and cheaply at scale. Deep reasoning is not the constraint here; throughput is. The correct move is the faster, lower-cost model, not the most capable one. Choosing Opus because it is "best" would burn budget and latency on drafts that never needed the horsepower; the volume is exactly what makes speed and cost dominate.
- Pick one representative task from your work (say, drafting a short reply to a routine customer message).
- Run it once on the fastest, lowest-cost model and once on the most capable one. Note roughly how long each took to come back.
- Compare the two outputs honestly. For a short, routine reply, is the more capable model's version meaningfully better, or just marginally different? Usually, for easy tasks, the gap is small, and you paid speed and cost for it.
- Now repeat the experiment with a genuinely hard task, a subtle analysis with competing considerations. Here the gap should open up, and the more capable model earns its cost. Those two experiments, side by side, are the trade-off this domain tests.
4. Context Limitations & Memory
A conversation with Claude has a working memory called the context window. Think of it as the amount of the conversation Claude can actively hold and reason over at once: everything you've said and it's said, plus anything you've attached. It is large, but it is not infinite, and a long, sprawling thread eventually pushes past what the model can keep in clear focus.
You don't get an error when that happens. You get drift. Learn to recognize the symptoms, because the fix depends on naming them:
- Details you established early quietly get dropped: a constraint you set twenty turns ago no longer holds.
- Claude contradicts something said far above, as if it never happened.
- Focus degrades: answers get vaguer, wander, or re-litigate things you'd already settled.
When you see those signs, you have three responses. Choosing the right one is the exam objective, not just knowing they exist.
| Response | When it's right | What you do |
|---|---|---|
| Restart | The thread is anchored to bad early turns, or the topic has genuinely changed and the old history is now noise. | Open a fresh chat and state the current goal cleanly, carrying only what still matters. |
| Summarize | You need continuity (the decisions and state so far) but not the verbatim back-and-forth that produced them. | Ask Claude to produce a condensed summary of where things stand, then continue from that compact state (often in a fresh thread). |
| Persist | The information is needed repeatedly, across sessions; it will outlive this conversation. | Move it out of chat entirely and into a Project's instructions or knowledge, where it applies to every future conversation instead of being lost. |
The distinction that trips people up is summarize versus persist. Summarizing keeps you moving within a piece of work: it's a within-task tool. Persisting is for information whose value is across work: if you'll need the same context next week and the week after, it doesn't belong in a chat you'll close, it belongs in a Project. Restart, by contrast, is the right call precisely when carrying history forward is the problem rather than the goal.
- Deliberately run a single chat long past where it should go: keep piling new sub-topics, tangents, and revisions into one thread without ever starting fresh.
- Watch for the first symptom of drift: a dropped constraint, a contradiction, a suddenly vaguer answer. Note what triggered it.
- Practice summarize-and-restart: ask Claude to summarize the decisions and current state in a compact form, open a new chat, paste that summary as the opening context, and continue. Notice focus return.
- Finally, identify one piece of context from that thread you'd need again next week. That's your persist candidate; put it into a Project's instructions instead of a chat, and confirm it now shows up automatically in a new conversation there.
5. Recognizing Which Situation You're In
Both decisions in this lab reduce to reading the situation correctly before you act. A few fast diagnostics:
- Surface: ask "will I need this again?" Recurring context → Project. Many sources to synthesize → research mode. A deliverable I'll keep refining → artifact. None of the above → plain chat.
- Model: ask "how hard is this task, really, and how much of it am I doing?" Easy and high-volume → the faster, cheaper model. Genuinely hard or high-stakes → the most capable one. Everything in between, which is most work → the balanced default.
- Context: ask "why is this thread degrading?" Anchored to bad early turns or off-topic now → restart. Need the state but not the transcript → summarize. Will need it across sessions → persist into a Project.
None of these require memorizing a spec sheet. They require honestly naming the shape of the work (its reuse, its difficulty, its volume) and letting that name pick the tool.
6. Lab Exercise: Match the Tool to the Work
Objective: make the two selection decisions deliberately on real work, and be able to justify each against the cost/speed/quality trade-off and the reuse test.
- Inventory. List four tasks you actually do: one throwaway question, one recurring task with standing context, one multi-source research question, and one deliverable you revise over time.
- Assign a surface. Map each to chat, Projects, research mode, or artifacts, and write one sentence saying why that surface, referencing persistence and reuse.
- Assign a model. For each task, pick Haiku, Sonnet, or Opus, and name which dial (cost, speed, or quality) drove the choice. If you wrote "Opus" more than once, challenge yourself: does that task truly need it, or is it habit?
- Run the high-volume case. Take the routine, repetitive task and run a batch of it on the faster, cheaper model. Confirm the quality holds for work this straightforward.
- Run the hard case. Take the genuinely complex task and run it on both the balanced and the most capable model. Decide whether the quality gain justified the added cost and latency, and record your reasoning.
- Manage a long thread. Take one of your ongoing conversations, drive it until it drifts, then apply the correct response (restart, summarize, or persist) and note which one the situation actually called for.
- Promote one thing to a Project. Find the piece of context you keep re-supplying and persist it into a Project's instructions. You'll build on this in Lab 5.
Step 7 is the bridge to the next domains: the difference between someone who chats with Claude and someone who builds durable, repeatable workflows is exactly this: knowing what to persist and where.
Check Yourself
Exam-style items. Commit to an answer before expanding.
A support team needs to generate a high volume of short, fairly routine customer-reply drafts every day. Turnaround speed and cost per draft matter far more than deep reasoning. Which choice best fits the task?
Correct: B. The task profile decides the model. Short, routine drafts at high volume don't demand heavy reasoning, so the constraints that dominate are speed and cost, exactly what the faster, lower-cost model optimizes. Matching the model to the task's real difficulty is the whole point of the trade-off.
A is the classic trap: "best model to be safe" spends cost and latency budget on quality the task never needed, and at volume that waste compounds. C misunderstands the problem: the surface's features aren't the cost driver, and disabling them degrades the work without addressing model fit. D throws away the whole toolset to answer a question that model selection already answers cleanly, and abandons everything else Claude does well.
An associate answers the same type of complex reporting request every week, and each time re-pastes the same company background, style guide, and reference figures into a fresh chat. What is the best way to work?
Correct: B. This is the textbook signal to persist: the same context is needed repeatedly, across sessions. A Project makes those instructions and that knowledge apply automatically to every conversation, so the associate stops re-supplying background and gets consistent output. The recurring, reusable nature of the work is what the Project surface exists for.
A still leaves you manually pasting every week: slightly faster friction is still friction, and nothing is persisted where Claude applies it automatically. C is exactly the anti-pattern from the context section: one endless thread drifts and drops details, and it's fragile, the opposite of durable. D is a category error: research mode is for multi-source synthesis, not for supplying standing context to a recurring task.
A strategy analysis has been running in a single chat for a long time. Claude has started contradicting a constraint set early on and its answers are getting vaguer. The analysis itself is still the right topic and the early decisions are still valid; the associate just needs to keep going with the current state intact. What is the best response?
Correct: B. The symptoms (a dropped constraint, contradictions, fading focus) say the thread has outgrown the context window. But the topic is unchanged and the early decisions are still good, so you need continuity without the verbatim transcript. Summarizing carries the state forward in a compact form and restores focus. That's precisely the case summarize is for.
A discards valid work: restart is for when the history is the problem (anchored to bad turns or a changed topic), which isn't the case here. C leaves the root cause (an overloaded context) untouched, so the drift only worsens. D swaps the model but keeps the same bloated conversation, so a more capable model still has to reason over a thread that has outgrown its window; the problem is the context, not the model.
Key Takeaways
- Pick the surface by reuse. Plain chat for one-off; Projects for recurring work needing standing context; research mode for multi-source synthesis with citations; artifacts for deliverables you iterate on and keep.
- Three model families, one trade-off. Haiku is fastest and cheapest for straightforward high-volume work; Sonnet is the balanced default for most professional work; Opus is most capable for genuinely complex reasoning, at higher cost and latency.
- The best model is not a safe default. Using Opus for everything wastes cost and speed; using the cheapest model for nuanced work produces output you have to redo. Match the model to the task's actual difficulty, and start from Sonnet, not the ceiling.
- The high-volume routine case wants the faster, cheaper model. When speed and cost dominate and reasoning is light, that is the correct choice, not the most capable one.
- Manage context deliberately. When a thread drifts, restart if it's anchored to bad turns or off-topic, summarize to keep state without the transcript, and persist into a Project when the information is needed across sessions.