Prompt Engineering Is Only Half the Job: Why AI Also Needs Response Rationalization
Updated: Sep 27
A conversation I had with a group of delegates while attending the International Lean Six Sigma Institute annual conference in Bratislava, Slovakia left me with a question I haven't been able to shake. As you might expect, artificial intelligence found its way into many conversations.
We discussed how AI will continue to change the "ecosystem" of Operational Excellence... the skills and competencies, the changing nature of work, and how organizations might build capability in a world where knowledge can increasingly be accessed at the moment of need. But one conversation went somewhere different.

The question was essentially this: "If organizations increasingly rely on AI for decision-support information involving process management, strategy, governance and more, how do we know that the responses guiding those decisions can be relied upon?"
Not whether AI is useful. Clearly, it is.
Not whether AI will become more capable. Clearly, it will.
The question was much more practical: "When an AI response may guide a decision, inspire action or influence behaviour, how do we establish sufficient confidence in the response before acting upon it?"
A Measurement-System Problem?
Coming from the world of quality and process improvement, my mind immediately went somewhere perhaps predictable... Measurement System Analysis (MSA).
Suppose several operators are using measuring devices to determine whether manufactured parts meet requirements. Before relying on the resulting data, we might conduct a Gauge R&R study. Or suppose several technicians visually inspect a product and make subjective judgments about whether particular defects are present. We might perform an Attribute Agreement Analysis.
Why? Because before making decisions based upon measurements, we want some confidence in the system producing those measurements.
Granted, the analogy isn't perfect, but it can help to frame the AI challenge we are all dealing with now and into the future. We are rapidly becoming capable of generating extraordinary quantities of analysis, recommendations and decision-support information. But what is our AI equivalent of Measurement System Analysis? How do we evaluate the reliability of the system producing the information upon which decisions may increasingly depend?
And here's where things become particularly interesting. Could We Simply Ask Another AI? AI operates at extraordinary speed. Humans don't. So perhaps the obvious solution is to have one AI review the response generated by another.
AI Number One provides the analysis... AI Number Two checks it... Problem solved! Perhaps not. Consider conventional inspection. If one inspector might miss a defect, adding another inspector may reduce some risk. But two inspectors agreeing doesn't necessarily mean the product is acceptable.
They may interpret the requirement the same incorrect way.
They may have received the same inadequate training.
They may use the same flawed inspection method.
They may both fail to recognize the same defect.

Agreement is not the same thing as accuracy. The same problem potentially exists when one AI evaluates another. Two AI systems might agree because the original answer is well supported. But they might also rely upon similar sources, assumptions, patterns or reasoning. Consensus can increase confidence but it doesn't automatically establish truth.
That led me to wonder whether we are approaching the problem from the wrong direction. Perhaps what we need isn't simply another inspector. Perhaps we need a discipline for deciding whether an AI-generated response provides sufficient reason to act.
I have recently been referring to that discipline as...
Response Rationalization
The term, however, requires some explanation... We often use rationalization negatively:
"To describe constructing reasons to justify something we have already decided to believe."
That is emphatically not what I mean here. The purpose of Response Rationalization is not to rationalize why an AI response must be correct. It is employed:
"To establish whether there are defensible reasons to rely upon it."
Put more simply... Look before you leap. And importantly, the objective isn't to prove why we shouldn't use an AI response. It is to determine whether we have sufficient reason to conclude that we can. That makes Response Rationalization the natural counterpart to another emerging AI discipline:
Prompt Engineering
Prompt Engineering Is Only Half the Job
Much has been written about Prompt Engineering. Ask better questions. Provide context. Define the desired outcome. Establish constraints. Refine the prompt.
All good advice, but Prompt Engineering primarily addresses the input: How do we ask AI better questions?
Response Rationalization addresses the output: Do we have sufficient reason to rely upon the answer?
A beautifully engineered prompt can still produce an answer that is wrong. More importantly, it can produce an answer that isn't obviously wrong. It may be:
Articulate.
Logical.
Detailed.
Persuasive.
"Mostly" correct.
Yet... it may still contain one assumption, omission or factual error significant enough to make acting upon it unwise. That may become one of the defining challenges of AI-enabled decision-making.
The Constraint Has Moved
For much of human history, obtaining information was difficult. Knowledge was scarce, expensive and frequently inaccessible. We travelled to libraries, consulted specialists, searched reference books and spent considerable time locating information before we could begin using it. The internet changed that. Generative AI has changed it again. Today, producing an answer may take seconds.
What has happened is the "constraint" has moved. Increasingly, the challenge isn't:
"Can I find an answer?"
It is:
"Can I determine whether this answer deserves my confidence?"
And perhaps that requires a different capability.
Five Questions Before You Leap
Response Rationalization doesn't need to become another bureaucratic approval process. In many situations, five questions may provide an effective starting point.

1. Can the important claims be independently verified?
Focus on the facts that materially affect the conclusion.
Can they be confirmed through reliable independent sources?
Can calculations be reproduced?
Do cited sources exist—and do they actually support what the AI says they support?
Not every sentence deserves investigation. Verify what matters.
2. Does the reasoning make sense?
Accurate facts can still produce an unsupported conclusion.
Follow the logic from evidence to recommendation.
Are steps missing?
Has correlation quietly become causation?
Has a general principle been applied to circumstances where it may not belong?
Think of this as inspecting the bridge between evidence and conclusion.
Both ends can look perfectly sound while something important is missing in the middle.
3. What assumptions is the response making?
Every recommendation contains assumptions. Some are stated. Many aren't.
Ask a deceptively simple question: "What would have to be true for this answer to be correct?" Then determine whether those conditions actually exist.
4. What might be missing?
This may be the most powerful question of all. AI responds to what we ask and to the information available to it. But important variables may never appear in the answer because they never appeared in the prompt.
What information could materially change the conclusion?
What perspective hasn't been considered?
What data would you want before making the decision yourself?
Sometimes the greatest weakness in an answer isn't something that is wrong. It is something that is absent.
5. What happens if the answer is wrong?
Not every AI response deserves the same scrutiny. Ask AI where you might have dinner tonight and a poor answer might result in a disappointing meal. Use AI to interpret a regulation, recommend a major investment, diagnose a manufacturing problem or support a safety-related decision and the consequences can be considerably greater.
Verification should therefore be proportional to risk. The greater the consequence of error, the greater the burden of validation.
From Trust to Proportional Confidence
That last point matters because Response Rationalization should not become an argument for distrusting AI. Quite the opposite... Its purpose is to help us use AI more confidently and responsibly.
The question isn't binary: Trust AI / Don't trust AI.
A better question is: How much confidence is justified given the evidence, uncertainty and consequence?
For low-risk applications, the answer may require almost no additional validation. For moderate-risk decisions, checking important claims and assumptions may be sufficient. For consequential, high risk decisions, we may require authoritative sources, independent calculations, human subject-matter expertise or other forms of validation.
Response Rationalization therefore introduces something valuable into an extraordinarily fast technology: Deliberate friction. Not enough to prevent us from moving. Just enough to make sure we look before we leap. Perhaps the emerging human-AI workflow is remarkably simple:

LEARN—That last step matters. It transforms our relationship with AI from a transaction into a learning cycle.
Perhaps AI Literacy Isn't Just About Using AI
Prompt Engineering will remain important. We should become better at communicating with increasingly capable AI systems no differently than effectively communicating with other human beings.
But something interesting may happen as those AI systems improve. AI may become increasingly good at compensating for our imperfect prompts. The greater challenge may eventually reside on the other side of the conversation...
Can we question an answer without automatically rejecting it?
Can we recognize uncertainty beneath confident language?
Can we distinguish evidence from inference?
Can we uncover assumptions?
Can we identify what is missing?
Can we determine when independent verification is warranted?
And ultimately:
Can we decide when there is sufficient reason to move forward?
Perhaps the future of AI literacy therefore requires two complementary disciplines:
Prompt Engineering helps us ask better questions.
Response Rationalization helps us decide whether we have good enough reasons to act on the answers.
Prompt Engineering improves the conversation. Response Rationalization improves the decision.
Answers are increasingly becoming almost effortless to obtain and knowing which answers deserve our confidence may prove to be one of the most valuable capabilities of all.
Continue Building Your AI Capability
Artificial intelligence is rapidly changing how we research, analyze, create, solve problems and make decisions.
The RPM-Academy Certificate in Artificial Intelligence explores AI concepts and practical applications designed to help individuals build the knowledge, competency and capability needed to work effectively in an AI-enabled world.




Comments