Why AI Context and the Right Model Matter When Building Apps
A tiny task, two very different answers
I gave two AI models the same three-line Python task. One gave me four lines of clean code and a useful warning. The other spotted the same edge case, but it quietly decided how to handle it for me.
The task was so simple that you'd expect both models to do it the same way. They didn't, and the difference taught me more than any benchmark. Context, and choosing the right model, matter even for the smallest tasks.
The task
I sent both models the same prompt, word for word:
- Build a Python script. You should define a function named
apply_discount. - The
apply_discountfunction should take exactly two parameters:priceanddiscount. - If
priceis not a number (int or float), the function should return the stringThe price should be a number.
That's the whole spec. It doesn't say how to handle the discount, what to return for valid input, or what to do about edge cases. A good assistant has to decide what to fill in on its own and what to ask me about.
Answer A: Claude Opus 5.5 kept it simple
Claude gave me four lines that do what I asked, and no more:
Claude Opus 5.5
def apply_discount(price, discount):
if not isinstance(price, (int, float)):
return 'The price should be a number'
return price - price * discount / 100
Three things stood out to me:
- It follows PEP 8, Python's official style guide. It uses a snake_case function name, 4-space indents, and spacing that any Python developer can read at a glance.
- It was open about its one assumption. My spec never said what the discount means, so Claude treated it as a percentage and said so, with an example:
apply_discount(100, 20)returns80.0. - It warned me about a real edge case, and then asked instead of guessing. In Python,
TrueandFalsecount as integers, soapply_discount(True, 10)would pass the number check. Claude explained this and asked whether I wanted to block booleans or add checks such as keeping the discount between 0 and 100. It didn't quietly build those features in.
The result is code I understand completely, plus a short list of decisions that are mine to make.
Answer B: Gemini 3.1 Pro decided for me
Gemini 3.1 Pro's code is about the same length, and it works:
Gemini 3.1 Pro
def apply_discount(price, discount):
# Check if the price is exactly an int or float, explicitly excluding booleans
# (since in Python, bool is a subclass of int)
if type(price) not in (int, float):
return "The price should be a number"
# Calculate and return the discounted price assuming discount is a decimal (e.g., 0.20 for 20%)
return price - (price * discount)
# Example usage:
print(apply_discount(100, 0.2)) # Output: 80.0
print(apply_discount("100", 0.2)) # Output: The price should be a number
print(apply_discount(True, 0.1)) # Output: The price should be a number
Gemini spotted the same boolean problem Claude did. The difference is what it did next: instead of asking me, it made the call itself.
- It changed the spec without asking. My requirement said "int or float", and
Truetechnically is an int. Blocking it is a design choice. Gemini made it for me, and only a code comment says so. - It made a different guess about the discount. Gemini treats it as a fraction (0.2), while Claude treats it as a percentage (20). Neither guess is wrong, but call Gemini's version with
apply_discount(100, 20)and you get-1900. An assumption nobody mentions turns into a real bug once a caller expects something else. - It broke a PEP 8 rule to do it. PEP 8 says type comparisons should use
isinstance()rather than checking the type directly. Gemini'stype(price) not in (int, float)does reject booleans, but it also rejects other number types, such as NumPy'sfloat64.
Neither answer is bad code. The difference is who made the decisions: Claude asked me, and Gemini decided for me.
Why context matters
An AI model only knows what you tell it. Anything you leave out, it has to guess, and every guess is a chance to build the wrong thing.
My three-line prompt left a lot unsaid:
- Is the discount a percentage (20) or a fraction (0.2)?
- What should happen when the discount is negative or above 100?
- Should
TrueandFalsecount as numbers? - Is this a learning exercise, a test you have to pass, or production code?
The last question matters most. A coding exercise wants exactly what the spec says, because extra behaviour can break hidden tests. A production app might need strict validation. Same prompt, different right answer, and only the context tells you which.
That's why the way a model handles missing context matters so much. A good model does three things: it solves what's clearly asked, says what it assumed, and asks about what's still open. A weaker one fills the gaps silently and leaves you to find out what it decided.
Why the model matters, even for simple tasks
It's easy to assume every top model handles a beginner task the same way. My test says otherwise. Both models are capable, but they have different habits, and on a vague prompt those habits decide what you get.
- Some models lean towards doing more. That helps when you want a full solution and hurts when you want exactly what you asked for.
- Some models lean towards asking. That keeps you in control, but it adds a round trip.
- No model is perfect. This was one small test on one task, not a benchmark. Run the same prompt again tomorrow and either model might answer differently, and on another task the roles could flip.
The lesson isn't "always use model X". Test the models you rely on with the kind of tasks you actually do, and pay attention to how they behave when your instructions are incomplete. That behaviour shows up in every project you build with them.
How to get better results
- Say what the code is for. "This is a freeCodeCamp exercise" and "this runs in our checkout service" call for very different code.
- State your limits. For example: "Keep it minimal", "Follow PEP 8", "Don't add validation I didn't ask for."
- Ask the model to list its assumptions. For example: "Before writing code, tell me anything that's unclear."
- Review the extras. Any behaviour you didn't ask for is a decision someone else made for you. Keep it or cut it on purpose.
- Compare models on your own work. Give two models the same real task and see which one behaves the way you want.
Conclusion
A three-line spec was enough to show the gap. One model gave me simple, readable code, said what it assumed, and asked about the boolean edge case. The other filled the gaps with its own guesses.
When you build apps with AI, the prompt is only half the job. The rest is giving the model the right context and picking a model that respects it. Do both and you'll spend your time building, not undoing decisions you never made.
Author: Rafał Tarłowski https://www.linkedin.com/in/rafal-tarlowski/
Comments
Post a Comment