Agentic AI changed what I have to prove before a question ships
I used to describe an AI pipeline by the thing it could produce. At Growtrics, that description became too small. The product needs educational content turned into structured questions and answers across subjects and curricula. An agent can generate a lot of them. The engineering question is which ones should be allowed into the product, which ones should wait, and what the whole system costs to run.
The public version of my work reports three figures: >99% precision, ~73% conversion of unstructured educational content into structured Q&A, and 50k+ questions populated. The portfolio also describes a structured-output schema, provenance-aware research and crawling, curriculum mapping, citation tracking, traces, and cost dashboards. Those are the reported facts I can share. The exact evaluation set, spend, and internal gate configuration are not public, so I will not invent them. This article is about the operating logic these numbers demand.
What changed for the developer
A conventional extraction function has a bounded contract: input document, output fields, or an error. An agentic workflow makes decisions along the way. It can search, choose a tool, inspect evidence, revise a draft, and stop. ReAct made that interleaving of reasoning and acting explicit in research [1]. The important shift for me is that every extra action creates another place where a plausible answer can lose its grounding.
That changes the developer's job. I still build parsers and interfaces, but I now also define the tool contract, source identity, acceptance criteria, abstention path, and evidence left behind for review. A fluent generated answer is a candidate, not a published item. Anthropic's distinction between a fixed workflow and an agent that chooses its own path is useful here [2]: add autonomy only where it earns its operational cost.
At Growtrics, this matters because a question can look perfectly formatted and still be pedagogically wrong: a statement attached to the wrong grade, a syllabus fragment mapped to the wrong chapter, or an explanation citing a page that does not support it. A schema catches shape; provenance and review have to catch meaning. The portfolio's stronger verification logic is the part I care about most, even if the headline number is the easiest thing to repeat.
The publication gate
The interactive model below is a design explanation, not a diagram of undisclosed Growtrics internals. A candidate question moves through four questions: Where did it come from? Is the structured record valid? Does the source support the claim? Can it be published, or should it be held? These checks may be implemented with deterministic rules, model-based reviewers, or human review depending on risk and cost. A reviewer is valuable when it has evidence to inspect and a real abstain option; asking the same model whether it feels confident is weaker.
This is why I resist one giant 'agent succeeded' metric. The output of the planner, parser, mapper, and reviewer should leave a trace of source IDs, decisions, and reasons for abstention. OpenAI's practical agent guidance recommends layered guardrails and human intervention for cases that need it [3]. The evaluator-optimizer pattern described by Anthropic is similarly useful when the evaluator has a crisp criterion, not just another opinion [2].
The '<1% FP' line needs a denominator
Precision is TP / (TP + FP) [4]. If the published >99% precision figure comes from a labeled evaluation of accepted questions, it implies FP / (TP + FP) <1% in that accepted set. This is also called the false discovery proportion. It is a meaningful quality statement: of what we let through, fewer than one in a hundred was judged wrong in that measurement.
It is not the same as the false-positive rate FP / (FP + TN), which asks what fraction of truly invalid candidates were mistakenly accepted. You cannot infer that rate from precision alone. Nor should a metric from one sample be quietly presented as a universal guarantee. To make a strong claim, I would publish the labeling rubric, sampling method, review period, denominator, and uncertainty interval. Those details are not in the public portfolio today. Even a clean audit needs scale: with zero errors in 300 independent, representative accepted items, the approximate one-sided 95% upper bound on the true error fraction is 3/300, or 1% [6]. That is an illustration of evidence strength, not a Growtrics audit result.
Precision is only useful next to conversion
A gate can get near-perfect precision by accepting almost nothing. That would not have served this product. The ~73% conversion figure says that a substantial share of unstructured educational material became structured Q&A. The 50k+ populated questions say the output was delivered at useful scale. Those metrics belong beside precision. They should not be multiplied together or treated as the same-cohort funnel without knowing their measurement windows and units.
The operating objective is a frontier: raise the acceptance threshold and published quality may improve while coverage falls and human review rises. Lower it and coverage may grow while false discoveries damage trust. The right setting depends on the harm of a wrong question, how quickly a learner or teacher can spot it, and what the review queue can absorb. An abstention is sometimes the correct product outcome, not a model failure.
The money has to sit on the same chart
A question is not 'cheap' because its model call was cheap. The cost numerator is crawler and storage spend, model calls across planning/extraction/review, human review time, and the infrastructure that keeps jobs traceable. The denominator is accepted, usable questions from the same period. Total cost / usable questions gives cost per question. The reported 50k+ figure counts populated questions. It does not establish how many passed a usable-quality audit in the exact period covered by a given spend figure, so I do not use it as the calculator denominator. Enter both period spend and same-period usable questions above to explore the ratio without pretending I have a public Growtrics invoice.
To compare two configurations, I would put four columns on one evaluation sheet: precision of published items, conversion or coverage, number of usable questions, and total cost per usable question. I would add review minutes and rework after publication. A higher precision number is not automatically better if it halves useful output or quietly doubles reviewer effort. A cheaper call is not cheaper if its errors return as manual cleanup.
What I would test before calling this strong
I would hold out documents by source and curriculum, not randomly shuffle near-duplicate pages between development and evaluation. I would label a representative sample of published items and a sample of rejected or held items. That second sample matters: it reveals whether the gate is too conservative and lets us estimate false-positive rate and missed useful questions. I would inspect slices where the system is most tempted to bluff: bad OCR, contradictory sources, edition changes, incomplete citations, and questions whose answer is technically true but wrong for the syllabus.
The metric should be attached to a versioned gate and a time window. Agent evaluations need realistic tasks and trace inspection, not just a benchmark score [5]. If the gate changes, I would remeasure both quality and conversion. If a reviewer starts rejecting everything, the precision dashboard may look wonderful while the product quietly stops growing.
The real change
Agentic AI did not remove engineering judgment. It moved judgment into the publication boundary. The work is to make every question earn its place, let uncertain material wait, and know the real price of the questions that survive. That is the version of agentic development I want to keep building.
References
- [1]Shunyu Yao et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models · ICLROriginal ReAct paper; reasoning/action loops, not the React UI framework.
- [2]Anthropic Engineering (2024). Building effective agents
- [3]OpenAI (2025). A practical guide to building agents
- [4]scikit-learn developers (2026). precision_score — scikit-learn documentation
- [5]Anthropic Engineering (2026). Demystifying evals for AI agents
- [6]Cochrane Handbook (2011). Confidence intervals when no events are observedRule of three for an approximate upper bound after zero observed errors.