Technology & Product Evaluation
This domain's work is mostly comparison under uncertainty: turning vague dissatisfaction into requirements, turning a noisy vendor market into a shortlist, and turning demos and claims into evidence a CFO will sign off on. Nearly every artifact is a structured document built from other people's persuasive documents — which is exactly where AI is both most useful and most dangerous.
The thesis · AI dramatically accelerates the *processing* of evaluation material (summarizing, scoring, cross-referencing, drafting test plans), but the *evidence-generating* steps — talking to users, running your own tests, checking references — must stay independent, because an AI summarizing a vendor's marketing inherits the vendor's framing.
Needs Assessment & Requirements Definition
per-selection (kicks off every evaluation; 2-6 weeks)stakeholder interview notes → pain-point inventory → prioritized requirements list (must/should/could) → evaluation criteria & weights
- DraftInterview prep — AI drafts role-specific interview guides from the project charter and org context · low stakes, easily verified by the analyst who knows the stakeholders, high repetition across roles
- AssistStakeholder interviews — humans conduct them; AI transcribes and produces per-interview summaries the interviewee confirms · the work is relationship-driven and the signal is in what people *don't* say; transcription is the automatable edge
- DraftPain-point synthesis — AI clusters themes across all transcripts and flags contradictions between departments · verifiable against the transcripts, medium stakes, and cross-document pattern-finding is where models beat a tired analyst
- DraftRequirements drafting — AI converts pain points into testable requirement statements ("system shall export to..." not "system should be user-friendly") · output is checkable line-by-line against the pain inventory; the human owns the must/should/could calls
- AvoidCriteria weighting — stakeholders negotiate what matters most; AI can only structure the trade-off discussion · this is a political consensus act, not an information task — weights set by AI have no organizational legitimacy and get relitigated later
- AutomateRequirements validation — AI checks the draft for untestable, duplicate, or vendor-biased requirements ("must use Salesforce" vs. "must sync with our CRM") · pure rules-vs-relationships rule-checking, fully verifiable, catches the classic failure of writing the answer into the question
Market Scan & Shortlisting
per-selection (1-3 weeks), plus an annual landscape refresh for owned categoriescategory landscape memo → long list (15-30 vendors) → screening matrix → shortlist (3-5) with rationale
- DraftLandscape mapping — AI with web search builds the long list and a first-pass feature/pricing table from vendor sites, G2, and analyst coverage · high volume and repetitive, but sources are marketing pages, so every cell is a claim to verify, not a fact
- AutomateClaim normalization — AI restates each vendor's self-description in neutral language and flags weasel phrasing ("AI-powered," "enterprise-grade," "seamless") · a deflationary rewrite is cheap, low-stakes, and instantly verifiable side-by-side with the original
- DraftViability screening — AI compiles funding, headcount trend, customer count, and acquisition-rumor signals per vendor · context is publicly available but frequently stale or wrong; a human spot-checks anything that would eliminate a vendor
- DraftScreening against must-haves — AI applies the requirements list to the long list and shows its evidence for each pass/fail · rule-like and verifiable *if* the evidence links are checked; vendor docs routinely overclaim, so "AI says it passes" is not a pass
- AssistPractitioner reality check — humans read Reddit/community threads, call peers at other firms, and ask "what do actual users complain about?" · the decisive signal lives in relationships and unindexed conversations; AI can summarize public threads but can't make the calls
- DraftShortlist memo — AI drafts the rationale memo from the screening matrix; the evaluator owns the cut lines · fully verifiable against the matrix, but the memo is a stakes-bearing commitment the analyst must be able to defend live
RFP & Vendor Response Evaluation
per-selection (for purchases large enough to warrant formal process; 4-8 weeks)RFP document → vendor responses → scored evaluation matrix → clarification questions → finalist recommendation
- DraftRFP drafting — AI assembles the RFP from the requirements list, prior RFPs, and standard T&Cs sections · templated and repetitive with verifiable inputs; legal and procurement still own the binding language
- AutomateResponse intake & normalization — AI extracts each vendor's answer to each question into a comparable grid, quoting the source passage · pure extraction with the quote as built-in verification; high volume, low judgment
- DraftFirst-pass scoring — AI scores each response against the published criteria with a quoted-evidence justification per score · verifiable because every score cites text, but vendors write to seduce scorers — human evaluators must re-score independently before seeing AI scores, or anchor bias contaminates the panel
- AutomateEvasion detection — AI flags answers that don't actually answer the question, hedge with "roadmap" language, or contradict another section of the same response · cross-reference work at volume that humans reliably skip; every flag is checkable in seconds
- DraftClarification round — AI drafts pointed follow-up questions from the flagged evasions; humans decide which to send · low stakes to draft, but question selection signals your priorities to vendors — a negotiation act
- AssistConsensus scoring session — the evaluation panel reconciles scores; AI's role is limited to surfacing where human scorers diverged most · the score of record must be defensible in a vendor protest; regulatory/procurement exposure means humans own every number
Proof-of-Concept & Pilot Design and Evaluation
per-selection for finalists (2-8 weeks per pilot)pilot charter (success criteria, sample, timebox) → test scenario scripts → execution log → results readout → go/no-go memo
- DraftSuccess criteria definition — AI stress-tests draft criteria for measurability and gaming ("adoption" vs. "weekly active use by the 12 pilot users on real work") · the critique is verifiable and the failure it prevents — an unfalsifiable pilot — is the most common one
- DraftTest scenario design — AI generates realistic task scripts and edge cases from the requirements and real workflow descriptions · high-volume generation a human curates; edge cases must be checked against how work actually happens, not how AI imagines it
- AvoidPilot execution — real users do real work in the tool; AI can log and lightly survey but must not perform the tasks · the entire point is generating independent evidence about human-tool fit; automating execution destroys the thing being measured
- DraftResults analysis — AI aggregates task success rates, time-on-task, and survey sentiment into a findings draft · verifiable against the execution log; watch for AI smoothing a messy "it depends" result into a clean narrative
- AssistAnomaly interrogation — humans interview the pilot users whose experience contradicted the averages · the pilot's highest-value signal is usually one user's dealbreaker; finding *why* is relationship work AI can only prep questions for
- DraftGo/no-go memo — AI drafts from the readout; the pilot owner signs it and presents it · high stakes and low reversibility once procurement proceeds, so the owner must be able to defend every line without the model in the room
Business Case & Selection Decision
per-selection (final 2-4 weeks), refreshed at renewalTCO model (3-5 year) → build-vs-buy-vs-assemble analysis → risk register → recommendation deck → signed decision memo
- DraftCost discovery — AI enumerates cost categories teams forget (integration labor, migration, training, seat growth, egress fees, exit costs) and drafts the TCO skeleton · checklist-like and verifiable, but every number must come from quotes and internal data, not model estimates
- DraftQuote and contract extraction — AI pulls pricing terms, escalators, auto-renewal clauses, and usage-tier cliffs out of vendor quotes and MSAs into the model · extraction is verifiable against the document, but a misread escalator clause compounds for years — spot-check every extracted number
- DraftBuild-vs-buy-vs-assemble framing — AI drafts the option analysis, including the newly viable "assemble" option (internal tools built with AI coding agents), with explicit maintenance-burden assumptions · genuinely useful structuring, but AI systematically underestimates the cost of maintaining what it helped build; a human must own the ongoing-ownership line
- AutomateSensitivity analysis — AI varies adoption rate, headcount growth, and price escalators and reports which assumptions flip the decision · mechanical, fully verifiable spreadsheet work with high repetition; the interesting output is which knob matters
- DraftRisk register — AI drafts vendor-viability, lock-in, and integration risks with mitigations, seeded from the scan-stage viability data · verifiable and low-cost to review; humans add the risks that come from institutional memory
- Avoid*Recommendation & decision — humans decide; AI's honest role is red-teaming the deck ("what would the strongest opponent of this choice say?") (Avoid for the decision itself) · high stakes, low reversibility, and accountability can't be delegated — the red-team pass is the Assist-shaped edge
Evaluating AI Products Specifically
per-selection, then quarterly re-evaluation (models and products churn too fast for one-and-done)golden task set (your real tasks + graded ideal outputs) → vendor claim audit → eval run results → security & data-handling review → adopt/monitor decision with re-test triggers
- AssistGolden set construction — humans collect 30-100 real examples of the task with known-good outputs; AI helps format and de-identify them · the golden set is the independent ground truth everything else rests on — generating it with AI would test the vendor against a model's imagination, not your work
- AutomateVendor claim audit — AI extracts every testable claim from the vendor's marketing and docs ("95% accuracy," "SOC 2," "no training on your data") into a claims-vs-evidence table · extraction with quoted sources, high volume, instantly verifiable — and the table's empty evidence column is the point
- DraftEval harness setup — AI drafts the scoring rubric and eval configuration to run the candidate against the golden set · templated work a non-engineer can verify by running it, but rubric wording quietly determines the results, so a human owns it
- DraftEval execution & scoring — the harness runs the product on your tasks; AI-as-judge does first-pass grading, humans grade a 20% sample to calibrate the judge · this is the recursion: an AI judging an AI needs the same audit you'd give a vendor claim, and the human-graded sample is what makes the judge's scores mean anything
- DraftSecurity & data-handling review — AI summarizes the vendor's DPA, retention terms, training-use policy, and subprocessor list against your checklist; security/legal validates · document analysis is verifiable, but regulatory exposure (FERPA, HIPAA, GDPR) means a professional signs off, and "we don't train on your data" requires reading the actual clause, not the FAQ
- AutomateChurn management — AI monitors vendor changelogs and model-version announcements and triggers a golden-set re-run when the underlying model changes · monitoring is repetitive and low-stakes; the re-run against your fixed golden set is the verifiable part — silent model swaps are now the norm, not the exception
- AvoidAdopt/monitor decision — humans weigh eval scores against cost, security posture, and switching risk · same as every selection decision: stakes and accountability sit with a person, and eval scores are one input, not the verdict
Source: Directing Intelligence course field guide, 2026. Tool lists are dated on purpose — they churn; the stage verdicts and their blockers are the durable part. Spot something the frontier has dissolved? Contribution is coming; for now, open an issue or PR on GitHub.