VivoLearn

Technology & Product Evaluation

This domain's work is mostly comparison under uncertainty: turning vague dissatisfaction into requirements, turning a noisy vendor market into a shortlist, and turning demos and claims into evidence a CFO will sign off on. Nearly every artifact is a structured document built from other people's persuasive documents — which is exactly where AI is both most useful and most dangerous.

The thesis · AI dramatically accelerates the *processing* of evaluation material (summarizing, scoring, cross-referencing, drafting test plans), but the *evidence-generating* steps — talking to users, running your own tests, checking references — must stay independent, because an AI summarizing a vendor's marketing inherits the vendor's framing.

Filter stages:

Needs Assessment & Requirements Definition

per-selection (kicks off every evaluation; 2-6 weeks)

stakeholder interview notes → pain-point inventory → prioritized requirements list (must/should/could) → evaluation criteria & weights

  • Draft
    Interview prepAI drafts role-specific interview guides from the project charter and org context · low stakes, easily verified by the analyst who knows the stakeholders, high repetition across roles
  • Assist
    Stakeholder interviewshumans conduct them; AI transcribes and produces per-interview summaries the interviewee confirms · the work is relationship-driven and the signal is in what people *don't* say; transcription is the automatable edge
  • Draft
    Pain-point synthesisAI clusters themes across all transcripts and flags contradictions between departments · verifiable against the transcripts, medium stakes, and cross-document pattern-finding is where models beat a tired analyst
  • Draft
    Requirements draftingAI converts pain points into testable requirement statements ("system shall export to..." not "system should be user-friendly") · output is checkable line-by-line against the pain inventory; the human owns the must/should/could calls
  • Avoid
    Criteria weightingstakeholders negotiate what matters most; AI can only structure the trade-off discussion · this is a political consensus act, not an information task — weights set by AI have no organizational legitimacy and get relitigated later
  • Automate
    Requirements validationAI checks the draft for untestable, duplicate, or vendor-biased requirements ("must use Salesforce" vs. "must sync with our CRM") · pure rules-vs-relationships rule-checking, fully verifiable, catches the classic failure of writing the answer into the question
Tools (2026)
Claude/ChatGPT with projects for synthesis, Granola or Fireflies for interview capture, Airtable for the requirements traceability matrix, Miro AI for affinity clustering in workshops
Failure mode
The team skips interviews and has AI generate a "typical requirements list for a CRM," producing generic requirements that every vendor satisfies on paper — so the evaluation can't discriminate.
Try it
Given four raw stakeholder interview transcripts for a university advising-software purchase, use AI to build a pain-point inventory and a 15-item testable requirements list, then manually flag which three requirements an AI invented that no stakeholder actually said.

Market Scan & Shortlisting

per-selection (1-3 weeks), plus an annual landscape refresh for owned categories

category landscape memo → long list (15-30 vendors) → screening matrix → shortlist (3-5) with rationale

  • Draft
    Landscape mappingAI with web search builds the long list and a first-pass feature/pricing table from vendor sites, G2, and analyst coverage · high volume and repetitive, but sources are marketing pages, so every cell is a claim to verify, not a fact
  • Automate
    Claim normalizationAI restates each vendor's self-description in neutral language and flags weasel phrasing ("AI-powered," "enterprise-grade," "seamless") · a deflationary rewrite is cheap, low-stakes, and instantly verifiable side-by-side with the original
  • Draft
    Viability screeningAI compiles funding, headcount trend, customer count, and acquisition-rumor signals per vendor · context is publicly available but frequently stale or wrong; a human spot-checks anything that would eliminate a vendor
  • Draft
    Screening against must-havesAI applies the requirements list to the long list and shows its evidence for each pass/fail · rule-like and verifiable *if* the evidence links are checked; vendor docs routinely overclaim, so "AI says it passes" is not a pass
  • Assist
    Practitioner reality checkhumans read Reddit/community threads, call peers at other firms, and ask "what do actual users complain about?" · the decisive signal lives in relationships and unindexed conversations; AI can summarize public threads but can't make the calls
  • Draft
    Shortlist memoAI drafts the rationale memo from the screening matrix; the evaluator owns the cut lines · fully verifiable against the matrix, but the memo is a stakes-bearing commitment the analyst must be able to defend live
Tools (2026)
Perplexity or ChatGPT/Claude with web search for the scan, G2 and Gartner Peer Insights as review sources, Vertice or Tropic for SaaS pricing benchmarks, a shared spreadsheet as the screening matrix of record
Failure mode
The AI-built comparison table silently treats vendor marketing claims as verified capabilities, and a vendor that "checks every box" on paper makes the shortlist over one that's actually better but describes itself modestly.
Try it
Pick a real category (e.g., survey tools), have AI generate a 12-vendor comparison table with cited sources, then audit 10 randomly chosen cells against the actual source pages and compute the table's error rate.

RFP & Vendor Response Evaluation

per-selection (for purchases large enough to warrant formal process; 4-8 weeks)

RFP document → vendor responses → scored evaluation matrix → clarification questions → finalist recommendation

  • Draft
    RFP draftingAI assembles the RFP from the requirements list, prior RFPs, and standard T&Cs sections · templated and repetitive with verifiable inputs; legal and procurement still own the binding language
  • Automate
    Response intake & normalizationAI extracts each vendor's answer to each question into a comparable grid, quoting the source passage · pure extraction with the quote as built-in verification; high volume, low judgment
  • Draft
    First-pass scoringAI scores each response against the published criteria with a quoted-evidence justification per score · verifiable because every score cites text, but vendors write to seduce scorers — human evaluators must re-score independently before seeing AI scores, or anchor bias contaminates the panel
  • Automate
    Evasion detectionAI flags answers that don't actually answer the question, hedge with "roadmap" language, or contradict another section of the same response · cross-reference work at volume that humans reliably skip; every flag is checkable in seconds
  • Draft
    Clarification roundAI drafts pointed follow-up questions from the flagged evasions; humans decide which to send · low stakes to draft, but question selection signals your priorities to vendors — a negotiation act
  • Assist
    Consensus scoring sessionthe evaluation panel reconciles scores; AI's role is limited to surfacing where human scorers diverged most · the score of record must be defensible in a vendor protest; regulatory/procurement exposure means humans own every number
Tools (2026)
Responsive or Loopio (many vendors now answer RFPs with these, which is itself worth flagging to students), Claude with long-context for 200-page response analysis, Notion or SharePoint as the evaluation workspace, spreadsheet scoring matrix of record
Failure mode
Evaluators read AI's scores before forming their own, and the "consensus" becomes a group of humans lightly editing a model's opinion of documents that were themselves AI-written — machine grading machine, with no independent judgment in the loop.
Try it
Give students two real-ish vendor responses to the same five RFP questions; half the class scores blind and half scores after seeing AI's scores, then compare the distributions to measure anchoring in their own room.

Proof-of-Concept & Pilot Design and Evaluation

per-selection for finalists (2-8 weeks per pilot)

pilot charter (success criteria, sample, timebox) → test scenario scripts → execution log → results readout → go/no-go memo

  • Draft
    Success criteria definitionAI stress-tests draft criteria for measurability and gaming ("adoption" vs. "weekly active use by the 12 pilot users on real work") · the critique is verifiable and the failure it prevents — an unfalsifiable pilot — is the most common one
  • Draft
    Test scenario designAI generates realistic task scripts and edge cases from the requirements and real workflow descriptions · high-volume generation a human curates; edge cases must be checked against how work actually happens, not how AI imagines it
  • Avoid
    Pilot executionreal users do real work in the tool; AI can log and lightly survey but must not perform the tasks · the entire point is generating independent evidence about human-tool fit; automating execution destroys the thing being measured
  • Draft
    Results analysisAI aggregates task success rates, time-on-task, and survey sentiment into a findings draft · verifiable against the execution log; watch for AI smoothing a messy "it depends" result into a clean narrative
  • Assist
    Anomaly interrogationhumans interview the pilot users whose experience contradicted the averages · the pilot's highest-value signal is usually one user's dealbreaker; finding *why* is relationship work AI can only prep questions for
  • Draft
    Go/no-go memoAI drafts from the readout; the pilot owner signs it and presents it · high stakes and low reversibility once procurement proceeds, so the owner must be able to defend every line without the model in the room
Tools (2026)
the candidate products themselves in sandbox tenants, Maze or UserTesting for structured task capture, Google Forms/Qualtrics plus AI analysis for pilot surveys, Claude/ChatGPT for scenario generation and log synthesis
Failure mode
The team lets AI generate plausible-looking test scenarios instead of deriving them from observed work, so the pilot validates that the tool handles imaginary tasks beautifully and misses the real workflow it breaks on day one.
Try it
Write a one-page pilot charter and ten test scenarios for evaluating a note-taking AI for student group projects, then trade with another team who must find the two scenarios that are unmeasurable or gameable.

Business Case & Selection Decision

per-selection (final 2-4 weeks), refreshed at renewal

TCO model (3-5 year) → build-vs-buy-vs-assemble analysis → risk register → recommendation deck → signed decision memo

  • Draft
    Cost discoveryAI enumerates cost categories teams forget (integration labor, migration, training, seat growth, egress fees, exit costs) and drafts the TCO skeleton · checklist-like and verifiable, but every number must come from quotes and internal data, not model estimates
  • Draft
    Quote and contract extractionAI pulls pricing terms, escalators, auto-renewal clauses, and usage-tier cliffs out of vendor quotes and MSAs into the model · extraction is verifiable against the document, but a misread escalator clause compounds for years — spot-check every extracted number
  • Draft
    Build-vs-buy-vs-assemble framingAI drafts the option analysis, including the newly viable "assemble" option (internal tools built with AI coding agents), with explicit maintenance-burden assumptions · genuinely useful structuring, but AI systematically underestimates the cost of maintaining what it helped build; a human must own the ongoing-ownership line
  • Automate
    Sensitivity analysisAI varies adoption rate, headcount growth, and price escalators and reports which assumptions flip the decision · mechanical, fully verifiable spreadsheet work with high repetition; the interesting output is which knob matters
  • Draft
    Risk registerAI drafts vendor-viability, lock-in, and integration risks with mitigations, seeded from the scan-stage viability data · verifiable and low-cost to review; humans add the risks that come from institutional memory
  • Avoid*
    Recommendation & decisionhumans decide; AI's honest role is red-teaming the deck ("what would the strongest opponent of this choice say?") (Avoid for the decision itself) · high stakes, low reversibility, and accountability can't be delegated — the red-team pass is the Assist-shaped edge
Tools (2026)
Excel/Google Sheets with Claude or ChatGPT working the model, Vertice/Tropic pricing benchmarks for negotiation leverage, Gamma or PowerPoint Copilot for the deck, the org's own finance templates as the format of record
Failure mode
AI fills TCO gaps with confident "typical industry" numbers that no one flags as fabricated, and the business case's precision (a 5-year NPV to the dollar) launders its fiction.
Try it
Build a 3-year TCO comparing buying a $12/user/month tool for 200 users vs. assembling an internal equivalent with AI coding tools, have AI run sensitivity on three assumptions, and identify which single assumption flips the answer.

Evaluating AI Products Specifically

per-selection, then quarterly re-evaluation (models and products churn too fast for one-and-done)

golden task set (your real tasks + graded ideal outputs) → vendor claim audit → eval run results → security & data-handling review → adopt/monitor decision with re-test triggers

  • Assist
    Golden set constructionhumans collect 30-100 real examples of the task with known-good outputs; AI helps format and de-identify them · the golden set is the independent ground truth everything else rests on — generating it with AI would test the vendor against a model's imagination, not your work
  • Automate
    Vendor claim auditAI extracts every testable claim from the vendor's marketing and docs ("95% accuracy," "SOC 2," "no training on your data") into a claims-vs-evidence table · extraction with quoted sources, high volume, instantly verifiable — and the table's empty evidence column is the point
  • Draft
    Eval harness setupAI drafts the scoring rubric and eval configuration to run the candidate against the golden set · templated work a non-engineer can verify by running it, but rubric wording quietly determines the results, so a human owns it
  • Draft
    Eval execution & scoringthe harness runs the product on your tasks; AI-as-judge does first-pass grading, humans grade a 20% sample to calibrate the judge · this is the recursion: an AI judging an AI needs the same audit you'd give a vendor claim, and the human-graded sample is what makes the judge's scores mean anything
  • Draft
    Security & data-handling reviewAI summarizes the vendor's DPA, retention terms, training-use policy, and subprocessor list against your checklist; security/legal validates · document analysis is verifiable, but regulatory exposure (FERPA, HIPAA, GDPR) means a professional signs off, and "we don't train on your data" requires reading the actual clause, not the FAQ
  • Automate
    Churn managementAI monitors vendor changelogs and model-version announcements and triggers a golden-set re-run when the underlying model changes · monitoring is repetitive and low-stakes; the re-run against your fixed golden set is the verifiable part — silent model swaps are now the norm, not the exception
  • Avoid
    Adopt/monitor decisionhumans weigh eval scores against cost, security posture, and switching risk · same as every selection decision: stakes and accountability sit with a person, and eval scores are one input, not the verdict
Tools (2026)
promptfoo or Braintrust for running evals without deep engineering, spreadsheet-based golden sets a non-engineer can maintain, OneTrust or Vanta-style trust-center pages for security artifacts, Claude/ChatGPT as calibrated first-pass judge, vendor changelogs plus an RSS/email watcher for churn triggers
Failure mode
The team evaluates the AI product by asking another AI whether it's good — demo impressions plus AI-summarized marketing — and never runs a single one of their own tasks through it, so they buy the best storyteller instead of the best performer.
Try it
In teams, build a 20-item golden set for one real task (e.g., summarizing case readings), run it through two AI products in their chat UIs, grade blind with a shared rubric, and compare your results to each vendor's accuracy claims.

Source: Directing Intelligence course field guide, 2026. Tool lists are dated on purpose — they churn; the stage verdicts and their blockers are the durable part. Spot something the frontier has dissolved? Contribution is coming; for now, open an issue or PR on GitHub.