When I began comparing AI tools across platforms, the goal was not to determine which one performed better, but to understand how each approached the same task. With the growing number of AI systems available, it is easy to assume they produce similar outputs when given the same prompt. In practice, that is not the case. Different platforms can generate noticeably different responses in terms of structure, completeness, tone, and even underlying assumptions. This raises an important question for organizations: if the same task produces different results depending on the tool used, how should those outputs be evaluated before being applied in a workplace setting?
This became especially relevant when I tested a prompt focused on developing a structured promotion and compensation framework within a municipal public safety system.
What Makes a Good Prompt
I have previously written about this in more detail (https://henrysuazo.wordpress.com/what-makes-a-good-prompt/), but it is worth briefly revisiting here. Strong prompts include: (1) a clearly defined task, (2) relevant contextual information, (3) specific criteria or constraints, and (4) an identified audience or tone.
These elements matter because AI systems rely on clarity and structure to produce aligned outputs. Without that structure, responses can become overly broad, inconsistent, or misaligned with the intended use case.
With this in mind, the following prompt was designed to reflect a realistic policy scenario:
Develop a structured promotion and compensation framework for personnel within a municipal public safety system that ensures pay is aligned with training, scope of practice, and clinical responsibility. Include: (1) A clearly defined progression hierarchy that distinguishes between Certified First Responders (CFR/firefighters), EMTs providing Basic Life Support (BLS), and paramedics providing Advanced Life Support (ALS), (2) An explanation of how increasing levels of training, certification, and clinical decision-making should correspond to increasing compensation, (3) Consideration of continuity of care responsibilities, including on-scene care versus patient transport and transfer to hospital care, (4) Identification of any inconsistencies in current role valuation and compensation structures, (5) A proposed model where compensation is proportionate to scope of practice, responsibility, and system demand. Write for a policy or union negotiation audience. Maintain a professional, evidence-based tone.
This prompt incorporated a defined task, clear context, specific constraints, and a target audience to ensure consistent responses across platforms.
I tested this prompt across Claude, Perplexity AI, and Gemini to evaluate differences in structure, clarity, completeness, and workplace fit. Each platform was assessed using the same prompt and a consistent set of criteria, including structure, completeness, alignment with requirements, and applicability to a workplace setting. While ChatGPT was not used in the comparison itself, it was helpful in refining the prompt prior to analysis.
What stood out early was that each platform demonstrated a different output style when responding to the same prompt. For example, Claude tended to organize its response into clearly defined sections with a logical progression, making it easy to follow and apply. Perplexity AI, in contrast, incorporated references and supporting information, which added credibility but occasionally disrupted the flow of the framework. Gemini often presented information in a more conversational format, which may be easier to read but less structured for direct implementation. These differences in output style have practical implications when applied to real organizational decisions.


One of the more important realizations that followed was how authoritative these systems can sound, even when the content has not been validated. There is a tendency to trust outputs that are well-organized or confidently written. Nevertheless, clarity does not always equal accuracy, and references do not automatically guarantee relevance. I also noticed instances where responses appeared influenced by factors beyond the immediate prompt. This suggests a risk of what I would describe as context drift, where adjacent or prior inputs may subtly influence the output. While not always obvious, this reinforces the need for verification.
No single platform consistently delivered a complete and reliable solution. Instead, each contributed something useful. Claude was effective for drafting structured content. Perplexity AI was better suited for validating claims and providing references. Gemini added value in shaping tone depending on the intended audience.
This led to a more practical approach. Rather than relying on one system, it is more effective to think in terms of a workflow: draft with one tool, validate with another, and then apply human judgment before using the output. This reduces the risk of over-reliance and encourages active evaluation.
If I were to approach this again, I would spend more time tightening the prompt and breaking the task into smaller steps. I would also be more deliberate about verifying outputs before comparing them.
For leaders, the takeaway is straightforward. While AI tools can accelerate thinking and provide useful starting points, they do not produce uniform results, even when given the same prompt. Differences in output style, structure, and underlying assumptions can influence how information is interpreted and applied. This places greater responsibility on the user to not only construct effective prompts, but also to actively evaluate and validate outputs before relying on them in decision-making. In that sense, AI is not a standalone solution, but a set of tools that require oversight, intentional use, and sound judgment.
