PITTECHS / ARTICLE

Your first AI project is probably not a chatbot.

Most teams open the AI conversation with should we add a chatbot. Sometimes the answer is yes. More often the first project that pays for itself is one specific workflow, picked on purpose, with a person still reviewing the output. Here is how to choose it.

Written for owners, founders and operations leads at companies of one to two hundred people who feel the pressure to adopt AI and do not yet have a plan.

Start here

Start with a workflow, not a tool

A chatbot is a shape, not a project. It can be the right shape once the knowledge base is clean, the permissions are clear and the answers are checkable. Before that, it is an open-ended system with no defined job, which is the hardest kind of software to evaluate and the easiest kind to abandon.

A workflow is different. Client intake, document review, email triage, follow-up, reporting, proposal drafting, an internal process that eats a day a week. A workflow has inputs, handoffs, decisions, tools and a review point. Those boundaries are what make the project scopeable: you can say what the system is allowed to do, what it must not touch, and where a person signs off.

That is also what makes it provable. A workflow has a before and an after, and somebody already knows what good looks like.

Candidate tests

What makes a workflow a good first AI candidate

The first builds that work tend to share six traits. A candidate missing two of them is usually a second project, not a first one.

The work happens every week, not every quarter.

AI is rarely worth building for a one-off annoyance. It earns its keep when the same shape of work arrives constantly: new inquiries, incoming documents, form submissions, support questions, quote requests, call notes, status updates. Repetition is what lets you define a good input, a good output, and a test that tells you the system still works next month.

The inputs are text, documents, email, forms or records.

Language models are useful around messy language work: summarizing, classifying, extracting, drafting, routing, and noticing what is missing. That is not the same as deciding. The useful version prepares the work so a person finishes it faster.

The outcome is something you already measure.

Faster response to a lead. Fewer missed follow-ups. Less retyping between two systems. Cleaner handoffs. If you cannot say what number moves, you cannot scope the build and you cannot tell afterwards whether it worked. Vague outcome, vague project.

There is an obvious place for a person to review the output.

The first build should assist, not replace. Keep a person on final client-facing messages, approvals, exceptions, and anything legal, medical, hiring, lending or insurance shaped. This is not timidity. It is what makes the project small enough to ship and safe enough to adopt.

Your current tools can stay where they are.

A good first build is usually a thin layer around the process you already run: read the form, summarize the thread, draft the reply, flag the missing document, write the field back to the CRM. Replacing the CRM is a different project with a different budget, and it is not the one that proves AI is worth doing here.

Somebody owns the process and will say whether the output is any good.

Not a sponsor. An owner, the person whose week gets better or worse. Without one, the project drifts into experimentation that nobody is willing to grade, and it quietly stops.

Scorecard

A scorecard for ranking the candidates

Pick three workflows. Score each one from 1 to 5 on the six dimensions below, then add them up. The point is not the arithmetic. The point is that it forces a comparison instead of a hunch, and it makes the reason for the choice legible to everyone who has to live with it.

Frequency

How often does this run? Daily beats weekly, weekly beats monthly. Score 5 for daily.

Time cost per cycle

How much human time does one pass consume? Look for reading, copying, reformatting, chasing and checking. Score 5 for an hour or more.

Handoff pain

Where does the work get delayed, dropped, duplicated or misunderstood? Work that crosses sales, admin, operations and leadership scores high.

Business impact

Would fixing this touch revenue, customer experience, staff capacity or what management can see? An internal convenience can be worth doing. It should not be first.

Data sensitivity

What would the system need to see? Score 5 when sanitized or internal examples are enough. Score low for protected health information, financial account data, employee records or production credentials. Low does not mean impossible. It means it needs written boundaries before anything moves, and it should not be the project you learn on.

Integration complexity

Can version one work from a form, a spreadsheet, an inbox or an export? Score 5 for that. Score 1 if five systems have to change on day one.

One rule about the total. A high score with no workflow owner is not a high score. If nobody will grade the output every week, move to the next candidate, whatever the number says.

Better first builds

Five first builds that beat a generic chatbot

Intake that stops being retyped.

Requests arrive as forms, email, PDFs and phone notes, then somebody retypes them into the system of record. Extract the fields, summarize the request, flag what is missing, and hand a person a filled draft to approve. The time saved is measurable from week one.

Document triage against a checklist.

For document-heavy teams, identify what arrived, compare it against the checklist the work actually requires, and flag the gaps. This is almost always more useful than chat with all our documents, because it answers a question somebody is already being paid to answer.

Routing for a shared inbox.

Classify the request, detect urgency, suggest an owner, draft the internal note. The assignment stays human. You remove the triage tax, not the judgment.

Follow-up support in the CRM.

The value leaks in the gaps: stale opportunities, missing notes, unstructured call summaries, follow-ups nobody sent. Draft the note, summarize the call, propose the next action, queue the update for review.

Retrieval over your own documents, scoped narrowly.

A question system over your SOPs, policies or support material can be a good first build when the scope is small and the answers are checkable. It has to cite the document and the passage, say it does not know rather than guess, and be tested against questions whose answers you already know. Those three conditions are the difference between a useful tool and a confident one.

What to avoid

Five mistakes that happen before the workflow is mapped

AI for everything.

A broad goal produces broad risk and nothing to point at in six weeks. Prove one process, then widen.

Connecting private data early to see what happens.

Customer records, patient data, financial account data, employee files, credentials. Start with workflow descriptions, sanitized examples and sample records. The boundary is cheaper to draw now than to walk back later.

Letting the model take the final decision in a high-risk flow.

Legal, medical, hiring, lending, insurance and compliance decisions. Software can prepare a decision. A person takes it. That line is worth holding on any budget.

Buying the tool before defining the process.

A subscription does not fix an unmapped workflow. If the owner, inputs, outputs, review step, data boundary and success criteria are unclear, the tool choice is premature by definition.

Judging it by the demo.

A demo proves the happy path on data you chose. Real use needs evals, error handling, access boundaries, cost awareness and a plan for the day the model is confidently wrong. The gap between those two is where most AI projects quietly die.

What a review produces

What a workflow review should produce

Whether you run this yourself or have somebody run it for you, the output should be decision-ready rather than a list of ideas. A review worth paying for ends with:

  • A map of the workflow as it actually runs, including the steps nobody documented.
  • The bottlenecks and the handoffs where work gets dropped.
  • The strongest two or three automation candidates, scored.
  • A data and privacy boundary: what the system may see, and what it must never.
  • One recommendation for what to build first, and why not the others yet.
  • What stays human-reviewed, written down before anyone builds.
  • What has to connect, and what can wait for version two.
  • What not to automate at all, with the reason.
  • A fixed scope and a price for the first build, if the opportunity is worth building.

Notice how much of that is subtraction. Most of the value in a first review is the two candidates it takes off the table and the one boundary it writes down, not the build it recommends.

Where to start

Where to start this week

Do not start with a tool list. Write down three workflows that are repetitive, expensive in hours and visible to the business. Score them on the six dimensions above. Name an owner for the winner. Then ask what the smallest useful version looks like, and whether it can ship in weeks rather than quarters.

The first AI project should be small enough to ship, important enough to matter, and bounded enough that the team trusts the output. That is usually not a generic chatbot. It is one real workflow, improved carefully.

Not sure which of the three to build?

Send us the workflows. We will map them, score them, write down the data boundary, and tell you which one to build first and why not the others yet. One written review, five business days from access. If the honest answer is that none of them is worth building this year, the review says that.

Start with a review

Reviews start at $500, fixed price agreed in writing before any work starts.
No secrets, credentials or confidential records through this form.