Methodology · updated 2026-09-14

Our Numbers and How We Count Them

This page defines the methodology behind every proprietary number Agentic Giants publishes — the 90% reduction in unsupported answers versus baseline vector RAG, 587+ engineering projects, SOC 2 Type II delivery, and the Guaranteed Pilot. Numbers without methodology are marketing. Numbers with a defensible test set, a named reviewer, and a real audit trail are what regulated buyers need before signing.

  • Every claim has a source. Each number below links to the concrete test set, count rule, or audit criterion it comes from.
  • Named engineers stand behind the method. Ahsan Ishfaq authored the measurement protocol; Ahmad Ishfaq reviewed the implementation. Both are on the record.
  • One number per claim, one page for method. Every other page across the site links back here. Nothing is quoted anywhere that this page does not defend.

The Marketing Number vs. The Method Behind It

Every row below is a claim that appears somewhere else on the site. Marketing takes the left column, method takes the right. Nothing goes in the left column without a right-column entry.

CapabilityMarketing NumberMethod Behind It (defensible in audit)
90% fewer unsupported answersvs baseline vector RAGInternal enterprise evaluation set; unsupported = fact not citable back to a graph entity; matched vector-only baseline held constant across the same 500 queries.
587+ engineering projects deliveredover 12+ yearsAggregate across the founding team's engineering history since 2013; a project counts when a shipped deliverable or invoiced engagement closed.
SOC 2 Type II deliveryfor regulated buyersSystems audited to SOC 2 Type II controls; attestation issued against your deployment by your auditor of record. Reference: RYVYL, SOC 2 Type II certified against systems we engineered.
6-week fixed-price pilotwith a written guaranteeTarget defined in writing before work begins (hallucination rate, latency, or workflow criterion). Miss the target inside 6 weeks and the pilot is free.

“We do not publish a hallucination rate we cannot reproduce on a client’s dataset in the first week of the pilot. The 90% figure is a floor, not a ceiling: it is the number we are comfortable defending against a matched baseline in every engagement we have shipped. If a competing measurement protocol produces a higher number for the same task, we would like to see it.”

Ahsan Ishfaq · AI Architect, Agentic Giants

Every Number, Every Method

How We Measured 90% Fewer Unsupported Answers vs Baseline Vector RAG

Unsupported answer rate is the fraction of model outputs whose factual claims cannot be traced back to a specific entity or relationship in the underlying knowledge graph. Our internal evaluation set covers 500 enterprise queries drawn from the fintech, healthcare, and SaaS engagements listed on /case-studies, with the client-identifying strings removed. Each query runs against two systems that share the same LLM and the same retrieval infrastructure below the retrieval boundary, isolating the graph vs. vector variable.

The baseline system is vector retrieval only (chunk similarity, no graph). The Agentic Giants system runs GraphRAG on the same corpus with the same model. Both systems output an answer plus a source list. A separate reviewer scores each answer as supported (every fact resolves to a source in the provided list) or unsupported (any fact fails to resolve). Across the 500 queries, the baseline system produces unsupported answers at roughly ten times the rate of the GraphRAG system — a 90% reduction. The measurement protocol was authored by Ahsan Ishfaq and reviewed by Ahmad Ishfaq before this page went live.

The figure is a floor because it excludes the pipeline hardening — the critic gate, human-in-the-loop approvals on state-changing tool calls, and MCP server scoping — that ships into every production engagement. See GraphRAG Implementation for the architecture that produces the number.

What Counts as “587+ Engineering Projects Delivered”

An engineering project counts when the founding team shipped software into production or delivered a signed-off engineering artifact — payment platforms, ledgers, mobile applications, data pipelines, machine-learning systems, knowledge graphs, and full-stack applications. Projects held in code review, prototypes that never crossed a production gate, and internal experiments do not count. The 587+ figure is the aggregate across the team’s 12+ years of engineering history, not engagements the Agentic Giants entity has invoiced since it was founded in 2024.

Public-name engagements documented on /case-studies — RYVYL, CaptureProof, Optevo, SAFER, mydiveo, Sufferfest, PACE Racing, REAP Pro, and others — represent the tip of that count. Named enterprise clients under NDA (a subset visible on the client logo strip) contribute the rest.

What “SOC 2 Type II Delivery” Means in Practice

SOC 2 Type II delivery is the engineering discipline of designing, deploying, and documenting systems so they can pass a SOC 2 Type II audit conducted against your deployment, by your auditor of record. The scope is defined by the American Institute of Certified Public Accountants’ Trust Services Criteria covering security, availability, processing integrity, confidentiality, and privacy — see the AICPA SOC 2 documentation.

A concrete reference: RYVYL is a SOC 2 Type II certified payments platform that Agentic Giants engineered end-to-end. The certificate is issued to RYVYL against RYVYL’s production systems, which is the correct legal posture for a delivery engagement. We do not sell a shared multi-tenant platform, so there is no organization-wide SOC 2 Type II certificate that applies to Agentic Giants itself.

The Guaranteed Pilot Standard

Every production pilot ships against a written target. Before work begins, we agree the acceptance criterion in a signed scoping document — a hallucination rate, a latency budget, a workflow success rate, or a compliance posture. The pilot window is typically six weeks. If we ship inside the window and hit the target, the pilot converts to a production engagement at the price agreed before work began. If we miss the target, the pilot is free.

The guarantee is enforceable because the target is defined against your workflow — not our template — and both parties sign it before code is written. The scoping document is the artifact your legal team will review; the pilot contract references it explicitly.

How We Track Engagement Extension Rate

Extension rate is the fraction of pilot engagements that convert into ongoing production support, ML operations, or a follow-on delivery contract inside 90 days of the pilot closing. The number is tracked in our CRM against the scoping document that opened the pilot and updated monthly. We report the rolling 12-month figure so a single quarter of unusual activity cannot skew the number that shows on other pages.

Frequently Asked Questions

If a proof question is not covered below, the fastest path is a briefing. We answer methodology questions on the phone the same way we answer them in an RFP.

Why do you publish methodology instead of just the marketing numbers?+

Every claim we make about production AI systems has to survive legal review, security review, and audit. Publishing the method behind each number is faster than answering the same question in every RFP. A prospect can read this page in five minutes and decide whether our proof is defensible before we ever get on a call.

What does the 90% reduction in unsupported answers actually measure?+

We measure the rate at which our GraphRAG systems produce answers whose facts cannot be cited back to an entity in the knowledge graph. On our internal enterprise evaluation set, that rate is 90% lower than a matched baseline built on vector-only retrieval. The methodology, dataset composition, and score definition are described in the section below.

What counts as an 'engineering project' in the 587+ figure?+

Any client engagement where the founding team shipped software into production or delivered a signed-off engineering deliverable — payment platforms, ledgers, mobile apps, data pipelines, ML systems, knowledge graphs. 587+ is the aggregate across the team's 12+ years of engineering history, not projects the Agentic Giants entity has invoiced since 2024.

Is Agentic Giants SOC 2 Type II certified, or SOC 2 Type II aligned?+

We deliver systems audited to SOC 2 Type II controls, with the attestation issued against your deployment by your auditor of record. Reference clients include RYVYL (SOC 2 Type II certified, real-time payments platform we engineered end-to-end), a healthcare product engagement, and REAP Pro (reappro.co — AI-powered real estate intelligence, SOC 2 Type II delivery in scope per client mandate). We do not sell a shared multi-tenant platform, so there is no company-wide SOC 2 Type II certificate for Agentic Giants itself.

What is the Guaranteed Pilot standard?+

Before a pilot starts, we agree the target — a hallucination rate, a latency budget, a workflow acceptance criterion — in writing. If we ship inside the agreed window (typically six weeks) and hit the target, the pilot converts to a production engagement. If we miss the target, the pilot is free.

About this page

This methodology page was written by Ahsan Ishfaq, AI Architect at Agentic Giants, and reviewed by Ahmad Ishfaq, AI Engineer. Every stat cited elsewhere on the site links here. When a number changes, this page updates first — and the dateModified timestamp in the Article schema updates with it.

Related: the AI Agent Development service page shows the architecture the 90% figure is measured on. The GraphRAG Implementation service page shows the retrieval layer that produces the reduction. The industry pages (fintech, healthcare) show the audit posture SOC 2 Type II delivery is built for.

Read next