AI Created a Job Aid: How Do I Test if it Works in Real Life?
After a decade in Learning & Development, I’ve learned one universal truth: the speed at which you can create content has almost zero correlation to the effectiveness of that content in the field. When GenAI hit the scene, I saw teams start churning out job aids at a rate that made my head spin. But speed isn't the metric that matters when your company is facing an audit—accuracy and utility are.
I keep a personal “hallucination log” on my desk—a physical notebook where I track the weird, confident mistakes AI makes when it’s trying to be helpful. It’s a sobering reminder that AI is a tool, not an instructional designer. If you’ve used AI to draft your latest check here performance support, you aren't done yet. In fact, the hardest part is just beginning. How do you move from a generated draft to a validated tool that actually survives the reality of a busy, messy workplace?

Before we add a single review step, we have to ask the most important question in the L&D toolkit: What is the risk if this is wrong?
The Risk-Based Validation FrameworkNot all job aids are created equal. A guide on how to update a password is a low-stakes task. A guide on handling hazardous chemical spills is a high-stakes task. If you apply the same level of scrutiny to both, you’ll burn out your SMEs. If you apply too little, you’re looking at a compliance nightmare.
I categorize every AI-generated asset using a simple risk matrix to determine the intensity of the performance support QA required.
Risk Level Content Type Validation Method Low Non-critical tips, software shortcuts, meeting etiquette. Self-review + peer edit. Medium Standard Operating Procedures (SOPs), process workflows. SME review + structured feedback. High Compliance, safety, legal, financial transactions. Workflow validation + 3rd party audit review. Fact-Checking AI: The Citation HabitAI is a probabilistic engine, not a truth engine. It likes to sound correct, which makes it particularly dangerous when it fabricates policies. When you’re vetting an AI draft, stop asking, “Does this sound right?” and start asking, “Where is the source?”
My rule of thumb: If the AI can’t point to the specific policy document, internal wiki, or legal memo that generated the step, I treat the step as a hallucination until proven otherwise.
When reviewing, I demand that the AI (or the designer) provide citations for every procedural step. If a step feels generic—like "click save"—verify it against the actual system UI. AI loves to hallucinate "Save" buttons that were moved or renamed in a software update three years ago.
The SME Review: Moving Beyond "Looks Good to Me"Nothing grinds my gears more than a blank email from an SME saying, "Looks good to me." That is not a review; that is an abdication of responsibility. If you want a job aid that works, you have to structure your SME review process to get actionable data.
Stop sending PDFs with "Please review." Instead, provide a structured feedback form for your job aid usability test. Ask specific questions:
Does this step reflect the current system constraints or only the “happy path”? Are there any “hidden” requirements (e.g., specific permissions) not mentioned here? If an employee follows this exactly, what is the most common way they would fail? Are there any regional or departmental nuances this guide ignores?By forcing the SME to look for potential failures rather than just grammatical errors, you engage their expertise far more effectively.
Workflow Validation: Testing in the WildThis is where most L&D teams fail. You can have the most beautiful, perfectly fact-checked document in the world, but if it doesn't fit into the workflow, it’s just digital clutter. This is where pilot testing comes in.
Do not skip this. Take your job aid to the floor. Find an employee who is not involved in the design process. Ask them to perform the task using only the job aid—no SME sign off template for eLearning help from you, no help from the SME.
The "Silent Observation" TechniqueDuring the pilot, sit back and watch. Do not offer guidance. If they get stuck, note where they get stuck. Do not jump in to save them. The moment you clarify a step, you have invalidated the test. If they can’t figure it out, the job aid is flawed, regardless of how "accurate" the steps are.
Workflow validation is essentially a usability test for your content. You are looking for:

I’ve kept my hallucination log for a reason. AI has specific "anti-patterns" that show up in job aids. You should be specifically looking for these during your QA phase:
The "Magic Step": The AI skips a crucial authentication step because it assumes you already have access. The "Or-Else" Fallacy: The AI suggests a path that bypasses a required compliance check or system field. The "Outdated Reference": The AI pulls from training data that includes deprecated software versions or old organizational structures. The "Passive Trap": Passive voice is the enemy of performance support. If the job aid says, “The report should be submitted,” replace it immediately with, “Submit the report.” Passive voice leads to confusion about who is responsible for the action. The "Performance Support QA" ChecklistBefore you ship anything—whether it was written by a human or an AI—run it through this final gauntlet. If you can’t check these boxes, don’t hit publish.
QA Task Status Named Owner assigned? (Who is responsible for the next update?) [ ] Direct, active voice throughout? (Eliminate passive voice.) [ ] Validated against the *actual* live system environment? [ ] SME provided written verification of technical accuracy? [ ] Pilot test conducted with at least one representative user? [ ] Version control and review cycle date established? [ ] Final Thoughts: Don't Overpromise on AIWe are currently in a hype cycle where people treat AI like an oracle. It isn't. It’s an intern who has read everything but understood nothing. Your job as an L&D practitioner is to be the filter. Your job is to be the skeptic. Your job is to ensure that when an employee pulls up that job aid in the middle of a stressful shift, they don't find a hallucination—they find a solution.
Ship content with a named owner. Write in the active voice. And for the love of everything, verify the steps yourself. If you’re not willing to test it in the wild, you’re not building performance support—you’re just adding to the noise.
Now, go look at your most recent AI-generated draft. What’s the risk if it’s wrong? If the answer is "more than zero," it’s time to stop editing and start testing.