Every AI Answer Looks Perfect Under Flat Light

A crew is finishing a corridor in a mid-rise, drywall hung and finished, painters scheduled for Monday. The temporary lights throw an even wash down the hallway and the walls look clean, flat, ready for the next trade. Then the finisher walks that same corridor with an inspection light held almost against the surface, sweeping it as they go. A long ridge appears down every butt joint, a row of shallow dimples marks a line of screws, and a cluster of pinholes sits where a coat of mud went on over an air pocket. The finish of the wall hasn't changed. The difference is the training and experience behind the critical eye of the person doing the inspection.

Steve Jost Profile Picture
Share

The tool always hands back a "finished" wall

Ask a chat tool for a subcontractor scope letter or a set of role descriptions, and what comes back is complete: full sentences, sensible headings, a confident tone, and no blank spot where it was unsure of itself. It looks like work somebody competent finished.

Sometimes it is. But these tools are built to produce the most plausible continuation of a request rather than a verified one; researchers writing in Nature found they will fluently make claims that are "both wrong and arbitrary." Plausible is exactly what a finished wall looks like under flat light.

Try it on a subject you know cold: ask a means and methods question in your own trade. The gaps will show up in the first paragraph, sometimes even in the first sentence. Ask the same tool something in a discipline you have never worked in, and the answer comes back with the same authority and the same clean formatting. Is the output well researched, means-tested, checked for inconsistencies, and well stated, or is it confidently incorrect and only formatted perfectly?

What the trained eye is doing

The finisher in that corridor is not simply looking harder than everyone else. They are checking against a specific list of things that go wrong in this type of work, and they know specifically where each one is because they've made the same errors.

The long edges of a drywall sheet come tapered from the factory, so joint compound sits in a shallow recess and finishes below the plane of the wall. Butt joints have no recess, so two cut ends meet flush and the compound must be feathered wide on both sides to make a slight crown read as flat. That is why butt joints telegraph through paint first.

One screw driven a hair proud will print through the paint for the life of the building; one driven too deep tears the face paper and gives up most of its holding power. Both stay invisible until the light hits them sideways.

The Gypsum Association defines finish levels 0 through 5 in GA-214: level 4 covers most walls taking flat paint, while level 5 adds a thin skim coat across the entire surface and gets called out where gloss paint or critical lighting will reveal everything. Level 4 versus level 5 matters most, because the finisher is not answering whether this is a good wall. They are answering whether this wall is right for this room, with this paint, under this light, on this job. A level 4 wall that passes in a warehouse fails in a lobby with a wall-washing fixture.

A real inspection needs three things: knowing the specific ways this work fails, a way to force those failures into view, and the standard that applies here. Someone who has never finished drywall has none of the three, so they walk the corridor, see a flat white wall, and sign it.

Nobody hands you a light for work you have never done

That is exactly where a company sits when it asks AI to produce something the company has never gotten right on its own.

A team that has fought change orders and implemented a proper change order process over twenty years can read an AI-drafted change order narrative and see in ten seconds when it misses basic information or key dates. The project management teams have built that skill and discernment over years. When nobody in the building has ever written a strong role description, an AI-drafted role description is a flat white wall.

The order that actually works

Do the work by hand enough times to have an opinion about it, and not once. Everything is too complex until you practice enough is the operating principle: the fourth pass shows you things that were invisible on the first.

Next, decide what "correct" looks like for you and your company; it will be different for every company, trade, and industry, and only you can make that call. Outside consultants such as D. Brown Management can help you come to that conclusion, but ultimately the decision rides with your team.

Then write it down and make it repeatable, because standards move through progressive levels of development, from something living in one person's head, to a document, to a process that gets trained and measured.

Then, finally, bring in the tool. By that point AI has something to work from, which is the same thing a good new hire gets: examples of finished work, the reasoning behind them, and a person qualified to grade the result. That order is people first, then process and tools, and running it backwards is how companies end up with expensive software wrapped around a process nobody ever defined.

Forty role descriptions and nobody's standard

A specialty contractor around $120 million in revenue has role descriptions on file for some of their positions, and they came from three directions. Some were written by the group manager and read like a list of everything the current person happens to do, some came from HR and are careful about compliance language while silent about the actual work, and a handful were written by new hires as a first-week exercise. All bad, in different ways and not covering all positions equally.

Leadership's instinct is reasonable: forty inconsistent or non-existent role definition documents, a tool that writes well, so put the pile in and ask for one clean set. What comes back is clean: consistent headings, parallel structure, plausible responsibilities under every title. The tool even suggests salaries, basic skills required for the roles, and some personality profiles for the position.

Then somebody must approve it, and the questions that make a role description useful in a contracting business are not the ones on the page. What decisions does this seat make without asking, and up to what dollar amount? Which of these tasks is the person expected to perform, and which are they expected to develop somebody else to perform? Job role complexity shifts with company size, stage of growth, and who occupies the seats nearby.

None of those absences look like errors, because they look like a clean document. So, the company approves output it cannot evaluate, and then hires, pays, promotes, and coaches against forty documents nobody in the building could defend. The wall got signed without anybody holding the light.

The better path costs more at the front and less overall. Leadership picks two or three roles they know cold, usually the ones they have hired for and been burned by and writes those by hand. They argue about them, which is where most of the value shows up, because the argument surfaces that two executives have different answers about who owns what part of the roles. Then they test the draft against the person sitting in the seat in one of the positions. What comes out is not three documents but a standard: the questions that matter at this company and the language the business already uses, which is where [selecting and managing people for a job role] must start, not a downloaded template.

Now AI has become useful. Hand it the three finished descriptions as reference material, give it the raw notes for the fourth, and ask for a draft that follows the pattern. The output can be graded, because two people in the room know what the good ones contain and why, and the tool has moved from writing the standard to applying it. An apprentice can read every manual ever written; inspecting the work still takes somebody with the field time.

Three times starting with AI is the right call

Learning. Asking the tool to digest and explain retainage rules in a new state is a reasonable use of five minutes. The output is not the deliverable; your understanding is, and it gets tested against a real source before it costs anything. One caveat: every tool has a knowledge cut-off, a date after which it knows nothing newer, so check how current its information is before leaning on it for research.

Practice. Use it as a sparring partner, having it argue the owner's position before a difficult negotiation, or asking it for ten weak role descriptions so a leadership team can say out loud why each falls short. Getting it wrong is free, and the repetition is the point.

Discussion. Print the AI's version of a role description, hand it to five people who do that work every day and let them tear it apart. The marked-up copy is worth more than the original, and the arguing builds the eye.

The test is what the output is for. In all three cases it feeds people who will judge it, and nothing is used as the work product. The trouble starts when the output is the deliverable and nobody present can grade it.

The "Raking Light Test"

Four questions, five minutes, in the meeting where somebody says, "why don't we just have AI do this."

1. Who here has done this work by hand, for this company, recently enough to argue about it? Name the person out loud, and if no name comes back, that is the finding, because better prompting will not fix it.

2. What would wrong look like on this deliverable? Write down three specific ways the output could be wrong while still reading clean; "it could have errors" does not count. Name real failures: the scope letter that drops the exclusions estimating always carries, or the role description that never states decision authority. A team that cannot name three does not have the light yet.

3. What does being wrong cost us? Put a number on the failures you just named. The scope letter that drops the exclusions costs the margin on that subcontract; the role description with no decision authority costs a mis-hire and a year of coaching. The bigger the number, the stronger the case for doing the reps first.

4. What standard are we checking against, and where is it written down? Level 4 or level 5. If the honest answer is that we would know it when we see it, the reps are the work and the tool comes second.

Looking ahead

The point here is sequence, not avoidance. AI is not magic and it is not a fad, and it is not the first technology this industry has bought in the wrong order. Grading the output takes firsthand knowledge of the work, not technical knowledge of AI, and every company already has that eye wherever it has done the reps.

Garbage in, garbage out has always been the rule with these tools and it still holds. What changed is that the garbage now comes back formatted, and a company without the eye to catch it will file it as a standard. So run the tool today where the eye already exists: meeting notes, first-draft correspondence, summarizing a long spec section, all fast to check. Where the eye does not exist yet, start small: one process, two or three reps done by hand, one written standard, then the tool. Most of the work D. Brown Management does sits in that gap between doing something well a few times and having it written down where it can be taught, and all relationships begin with a conversation.



Related Training

Definition - Delegation
Delegation is the process of distributing and entrusting work to another person. This includes standards, resource allocation, follow-up, and quality checks to ensure the work is done correctly. Accountability for outcomes does not diminish.
Five Mountains - Where Are You At?
As you think about your career and business, there are five basic mountains to climb. Each has its own challenges and rewards. Each has its own peaks and valleys during the climb. Each progressively adds more value to the next generation.
Management Team Development Phases Within Each Stage of Growth
The management team leading a contractor through the different stages of growth will typically navigate three phases at each stage - Emerging, Hollow, and Ready. Understanding these phases guides growth planning, recruitment, development, and succession.