Skip to content

Accuracy

How accurate is FileAI?

We measure the checks below on real public documents against an answer key, and we publish the misses. Pick a check to see its numbers.

How accurate is the tender breakdown?

We ran the tender pack breakdown on 12 public Australian tender packs (521 pages) and checked every row against an answer key. Here is what it found, what it missed, and exactly how we measured it.

Last run 6 October 2026 · 12 packs from NSW, SA, TAS, VIC, WA · measured with GPT-6.1 Sol

Must-haves found
98% 731 of 746 things to do before you submit
Quotes on the cited page
99.7% 3,386 of 3,395 rows, re-checked independently
Rows that are real requirements
99% Estimate from a sample of 50 checked rows
Time per pack
2.6 min AI reading time for an average 43-page pack. Your free preview usually arrives within 15 minutes, longer for big packs.

The bar we set before publishing: find at least 90% of must-haves with at least 95% of quotes correct. This run meets it.

FileAI compared with one big prompt

We also gave the same AI model each whole pack in a single request with a one-line instruction: “list everything a bidder must do or provide, with page numbers and quotes”. That is roughly what you get by pasting a tender into a chatbot. Both were scored the same way.

Measure FileAI Single prompt
Must-haves found before you submit 98% 731 of 746 91.8% 685 of 746
Contract obligations found if you win 98.2% 380 of 387 93.5% 362 of 387
Quotes found on the cited page 99.7% 3,386 of 3,395 99.8% 1,973 of 1,977
Rows that are real requirements estimate, 95% range 99% range 94.5–99.8% 95.9% range 90.4–98.4%
Rows per pack 282.9 164.8
Pages per pack 43.4 43.4
Minutes per pack 2.6 3.3

The single prompt found most of the forms and documents to return, but only half of the evaluation criteria and about four in five key dates, and it listed a little over half as many rows. FileAI reads each pack in sections of about eight pages with written rules for every type of requirement, then merges the rows and checks every quote against its page.

Where it misses

Must-haves found by type. The misses are listed honestly below, so you know what to double-check yourself.

Type FileAI Single prompt
Condition for participation 100% 25 of 25 100% 25 of 25
Mandatory requirement 100% 77 of 77 98.7% 76 of 77
Returnable document 97.7% 338 of 346 99.4% 344 of 346
Insurance 95.7% 22 of 23 100% 23 of 23
Licence or certification 94.7% 18 of 19 100% 19 of 19
Key date 100% 28 of 28 78.6% 22 of 28
Evaluation criterion 98.7% 76 of 77 49.4% 38 of 77
Submission rule 97.4% 147 of 151 91.4% 138 of 151

Contract obligations found (if you win)

Type FileAI Single prompt
Condition for participation 100% 1 of 1 100% 1 of 1
Mandatory requirement 97.8% 270 of 276 96% 265 of 276
Insurance 98% 49 of 50 98% 49 of 50
Licence or certification 100% 20 of 20 100% 20 of 20
Key date 100% 40 of 40 67.5% 27 of 40

Single lines in long price schedules

When a schedule lists dozens of rates, FileAI can summarise part of it and drop a line, such as a site establishment rate, one trade in a list of four, or one size of pavement numeral. Check the schedule itself before you price.

A second instruction in the same sentence

Where one sentence asks for two things, like signing the offer and answering every selection criterion, or giving certificate details and attaching a copy, the second part is occasionally lost.

Criteria listed twice

One pack listed what it would consider in two places with different wording. FileAI captured the weighted table and only part of the second, unweighted list.

Contract details we summarise on purpose

Cleaning checklists and similar schedules get one row per area, not per task, so a once-a-year task is only in the schedule itself. The other contract misses were a payment reduction scale, a no-claims clause, an intellectual property term and part of an auditor's reporting duty.

Pages that are pictures

Scanned pages, drawings and tables saved as images have no text to read, so they are outside these numbers. FileAI names them as pages it couldn't read so you can check them yourself.

What we tested

12 real tender and quotation packs (14 PDF files, 521 pages) published on official council and state agency websites in NSW, SA, TAS, VIC, WA. They cover cleaning, IT, building and civil works, planning and audit consultancies, marketing and a supply panel, from 9 to 111 pages each.

The packs are used for testing only. We don't republish them, so they are described by state, buyer and work type, not by name.

  • WA local government · Facilities / cleaning services

    79 pages · 92 items in the answer key

    FileAI found 97.8%, single prompt 94.6%

  • WA local government · Professional services

    51 pages · 56 items in the answer key

    FileAI found 96.4%, single prompt 100%

  • WA local government · IT managed services / cybersecurity

    23 pages · 81 items in the answer key

    FileAI found 97.5%, single prompt 91.4%

  • WA local government · Construction

    44 pages · 66 items in the answer key

    FileAI found 98.5%, single prompt 93.9%

  • WA local government · Professional services

    43 pages · 85 items in the answer key

    FileAI found 98.8%, single prompt 95.3%

  • TAS local government · Professional services

    41 pages · 33 items in the answer key

    FileAI found 100%, single prompt 69.7%

  • SA local government · Professional services

    22 pages · 57 items in the answer key

    FileAI found 98.3%, single prompt 87.7%

  • NSW local government · Construction

    30 pages · 28 items in the answer key

    FileAI found 96.4%, single prompt 82.1%

  • NSW local government · Council works and services panel

    43 pages · 50 items in the answer key

    FileAI found 94%, single prompt 88%

  • NSW local government · Civil services

    9 pages · 35 items in the answer key

    FileAI found 97.1%, single prompt 94.3%

  • VIC state agency · Construction

    111 pages · 86 items in the answer key

    FileAI found 98.8%, single prompt 89.5%

  • VIC state statutory body · Marketing / communications services

    25 pages · 77 items in the answer key

    FileAI found 100%, single prompt 97.4%

How we measured it

The answer key

An AI review pass (a different AI model from the one FileAI uses) read every page and listed every obligation: 1,133 in all. Each entry's quote was then checked by a program against the text of its page.

No person has reviewed the whole key. It will have its own misses and judgement calls, so treat these numbers as a careful estimate, not a guarantee.

Must-haves before you submit (746) were listed exhaustively. Contract obligations if you win (387) were listed for the material items only, such as insurances, licences, plans and key dates.

When an item counts as found

A row has to cite the same file and a page within one page of the answer key's page, and say the same thing. A grading model decides the meaning with strict rules: a related but different item, a row that leaves out the deadline or amount, or a different number does not count.

We checked the grader by hand: on 40 decisions we reviewed ourselves, it agreed with us 37 times (92.5%).

An earlier version of the grading rules was too strict: on 40 other decisions it agreed with us only 28 times, always by marking a found item as missed. We fixed the rules, then ran the check above on new decisions.

When a quote counts as correct

The quote must appear on the page it cites, word for word, ignoring only spacing, capital letters and the style of quote marks and dashes. A quote may start on the cited page and run onto the next. Rows from tables count when every word of the quote sits together in that part of the page (21 of FileAI's 3,395 rows). This check is separate from the one FileAI runs on itself.

Extra rows and time

FileAI lists more rows than the key, because the key is selective: it holds the material items, not every line. To check those extra rows are real, we graded a random sample of 50 of them against their page.

Time is the AI processing time for a pack, with sections read in parallel the way the service runs them. It leaves out waiting in the queue and sending the email, and the person who checks every row, so a breakdown takes longer than this to reach you: after you upload, the free preview is usually ready within 15 minutes, longer for big packs, and the checked full breakdown usually within one business day.

The grader read 50 sampled FileAI rows that matched nothing in the answer key and found 49 of them to be real requirements. The one it rejected turns something the buyer may do into a task for you. We hand-checked the grader on an earlier run (see above), not on this sample.

Limits of this test

  • 12 packs is a small sample, and there are none from Queensland, the NT or the ACT.
  • The answer key is AI-made and not fully human-reviewed.
  • AI output varies a little from run to run; this is one run of each system.
  • We adjusted FileAI's instructions using the misses on these same packs (on an earlier model), so expect slightly lower results on a tender it hasn't seen. On five packs that an earlier version had never been tuned on, it found 97% of must-haves.
  • Pages that are scanned images or drawings have no text we can read. FileAI lists them as pages it couldn't read, so you can check them yourself.

A person also checks every row

While volumes are low, a person reviews every row of a paid breakdown against its page before it is delivered. Every row still links to its page so you can check it yourself.

How accurate is the RFP compliance matrix?

We ran the RFP compliance matrix on 12 real public US solicitations (952 pages) from federal, state and local buyers, and checked every row against an answer key written from the documents themselves. Here is what it found, what it missed, and exactly how we measured it.

Last run 6 October 2026 · 12 solicitations: CA, Federal, MO, MS, NY, TN, TX · measured with GPT-6.1 Sol

Proposal requirements found
96.4% 1,603 of 1,663 things to include, do or meet in your proposal
Quotes on the cited page
99.1% 7,551 of 7,623 rows, re-checked independently
Rows that are real requirements
91.3% Estimate from a sample of 50 checked rows
Time per RFP
4.9 min AI reading time for an average 79-page RFP. Your free preview usually arrives within 15 minutes, longer for big RFPs.

The bar we set before publishing: find at least 90% of proposal requirements with at least 95% of quotes correct. This run meets it.

FileAI compared with one big prompt

We also gave the same AI model each whole solicitation in a single request with a one-line instruction: “list everything an offeror must do or provide, with page numbers and quotes”. That is roughly what you get by pasting an RFP into a chatbot. Both were scored the same way.

Measure FileAI Single prompt
Proposal requirements found for your proposal 96.4% 1,603 of 1,663 68.5% 1,140 of 1,663
Contract terms found if you win 95.5% 621 of 650 64.5% 419 of 650
Quotes found on the cited page 99.1% 7,551 of 7,623 99.7% 3,297 of 3,308
Rows that are real requirements estimate, 95% range 91.3% range 84.5–95.5% 98.2% range 93.9–99.5%
Rows per RFP 635.3 275.7
Pages per RFP 79.3 79.3
Minutes per RFP 4.9 4.7

Where it misses

Proposal requirements found by type. The misses are listed honestly below, so you know what to double-check yourself.

Type FileAI Single prompt
Submission instruction 97.2% 411 of 423 79.9% 338 of 423
Evaluation factor 96.6% 228 of 236 11.9% 28 of 236
Required form or certification 94.5% 191 of 202 78.7% 159 of 202
Eligibility and registration 96% 24 of 25 72% 18 of 25
Key date 95.1% 117 of 123 52.8% 65 of 123
Mandatory requirement 97.5% 426 of 437 85.8% 375 of 437
Past performance or personnel requirement 97% 97 of 100 85% 85 of 100
Pricing instruction 93.2% 109 of 117 61.5% 72 of 117

Contract terms found (if you win)

Type FileAI Single prompt
Key date 75% 3 of 4 75% 3 of 4
Mandatory requirement 100% 6 of 6 100% 6 of 6
Pricing instruction 100% 1 of 1 100% 1 of 1
Clause of note 95.3% 407 of 427 60.9% 260 of 427
Insurance or bonding requirement 96.4% 106 of 110 80.9% 89 of 110
Reporting or deliverable 96.1% 49 of 51 72.5% 37 of 51
Payment and invoicing 96.1% 49 of 51 45.1% 23 of 51

What an amendment changed

When an amendment replaces a price workbook, adds a line item or repeats a revised calendar, FileAI lists the new deadline but can leave out what else the amendment changed. Read every amendment's cover page against your matrix.

Single clauses in long flow-down lists

Federal and grant-funded contracts (FTA, CDBG) list dozens of clauses. FileAI gives each clause that carries a duty its own row, but can still drop one, such as a change-order rule or a facility security requirement. Check the clause list yourself before you sign.

Representations by number

Section K representations that appear only as a clause number and title (for example covered telecommunications equipment, or delinquent tax and felony convictions) are sometimes missed. Complete Section K from the solicitation itself.

Price workbook detail

Instructions inside a pricing schedule, such as how option periods or an extension period are priced, or milestone payment percentages, are the proposal group FileAI misses most often after evaluation factors. Read the price instructions and the workbook together.

Go/no-go gates in multi-phase evaluations

Where a solicitation scores in phases with pass/fail gates, a gate's exact criteria (minimum project values, what counts as similar work) can be shortened or missed. Treat every go/no-go statement as a must-have.

What we tested

12 real public solicitations (39 PDF files, 952 pages) published on official federal, state and local government procurement sites. They cover services and construction, from 28 to 141 pages each.

The solicitations are used for testing only. We don't republish them, so they are described by buyer level and work type, not by name.

Tuned sets and held-out sets

We improved FileAI's instructions while looking at 6 of the 12 sets (marked in the table). The other 6 were not used to make those changes and were first scored in the final run, so they show how it does on documents it was not tuned on.

Sets Proposal requirements found Contract terms found
Used for tuning 6 sets, 310 pages 97.9% 546 of 558 97.4% 294 of 302
Held out 6 sets, 642 pages 95.7% 1,057 of 1,105 94% 327 of 348
  • CA city · Construction (renovation) (used for tuning)

    52 pages · 73 items in the answer key

    FileAI found 97.3%, single prompt 83.6%

  • Federal agency · Medical services

    125 pages · 172 items in the answer key

    FileAI found 100%, single prompt 66.3%

  • Federal agency · Transport services (RFQ) (used for tuning)

    28 pages · 63 items in the answer key

    FileAI found 98.4%, single prompt 73%

  • Federal agency · Research services

    87 pages · 218 items in the answer key

    FileAI found 93.1%, single prompt 59.6%

  • Federal agency · Healthcare services

    136 pages · 232 items in the answer key

    FileAI found 95.3%, single prompt 73.7%

  • Federal agency · Construction (used for tuning)

    45 pages · 77 items in the answer key

    FileAI found 98.7%, single prompt 0%

  • Federal agency · Design-build construction (used for tuning)

    83 pages · 185 items in the answer key

    FileAI found 97.3%, single prompt 66%

  • MO state agency · Construction (used for tuning)

    65 pages · 71 items in the answer key

    FileAI found 97.2%, single prompt 63.4%

  • MS state agency · IT services

    141 pages · 179 items in the answer key

    FileAI found 96.1%, single prompt 84.9%

  • NY state agency · Training services

    74 pages · 180 items in the answer key

    FileAI found 95.6%, single prompt 70%

  • TN public authority · Professional services

    79 pages · 124 items in the answer key

    FileAI found 94.3%, single prompt 79.8%

  • TX city · Planning consulting (used for tuning)

    37 pages · 89 items in the answer key

    FileAI found 98.9%, single prompt 83.2%

How we measured it

The answer key

The answer key was written by reading every solicitation page by page and listing what a proposal manager's compliance matrix must contain: 2,313 rows in all. It was written before FileAI was run on the documents, by an AI assistant following written rules, not by FileAI. A program then checked that every quote is on the page it cites.

No procurement professional has reviewed the key. It will have its own misses and judgement calls, so treat these numbers as a careful estimate, not a guarantee.

Proposal requirements (1,663) were listed exhaustively: submission instructions, evaluation factors, required forms and certifications, key dates, mandatory requirements of the work, past performance and key personnel, pricing instructions. Contract terms if you win (650) were listed for the material items only, such as clauses of note, insurance and bonding.

When an item counts as found

A row has to cite the same file and a page within one page of the answer key's page, and say the same thing. A grading model decides the meaning with strict rules: a related but different item, a row that leaves out the deadline or amount, or a different number does not count.

We have not yet checked the grader by hand on RFP rows. It uses the same strict rules as the tender grader, which agreed with our own labels on 37 of 40 tender decisions, so treat the RFP recall as a careful estimate. Nine answer-key rows got no verdict from the grader and are counted as missed.

When a quote counts as correct

The quote must appear on the page it cites, word for word, ignoring only spacing, capital letters and the style of quote marks and dashes. A quote may start on the cited page and run onto the next. Rows from tables count when every word of the quote sits together in that part of the page (283 of FileAI's 7,623 rows). This check is separate from the one FileAI runs on itself.

Extra rows and time

FileAI lists more rows than the key, because the key is selective: it holds the material items, not every line. To check those extra rows are real, we graded a random sample of 50 of them against their page.

Time is the AI processing time for a RFP, with sections read in parallel the way the service runs them. It leaves out waiting in the queue and sending the email, and the person who checks every row, so a breakdown takes longer than this to reach you: after you upload, the free preview is usually ready within 15 minutes, longer for big RFPs, and the checked full breakdown usually within one business day.

We graded a random sample of 50 of the rows that matched nothing in the answer key: 42 were real requirements, 5 described something the buyer does rather than a task for you, and 3 quoted text the grader could not find on the cited page.

Limits of this test

  • 12 solicitations is a small sample of the thousands published each year.
  • The answer key was written by an AI assistant and not reviewed by a person.
  • AI output varies a little from run to run; this is one run of each system.
  • We revised FileAI's RFP instructions after reading the misses of a first run, and measured them on 6 of the 12 solicitations while revising. The other 6 were scored once, at the end: on those it found 95.7% of proposal requirements and 94.0% of contract terms, against 97.9% and 97.4% on the 6 used for tuning. The instructions contain no wording or figures from any of the 12.
  • Pages that are scanned images or drawings have no text we can read. FileAI lists them as pages it couldn't read, so you can check them yourself.

A person also checks every row

While volumes are low, a person reviews every row of a paid breakdown against its page before it is delivered. Every row still links to its page so you can check it yourself.

How accurate is the HOA document check?

We ran the HOA document check on 6 real sets of public HOA and condominium association documents (632 pages) and checked every row against an answer key written from the documents themselves. Here is what it found, what it missed, and exactly how we measured it.

Last run 6 October 2026 · 6 document sets: AZ, CA, DE, FL, VA, WA · measured with GPT-6.1 Sol

Findings found
91% 193 of 212 fees, assessments, reserves, lawsuits, insurance, rules and repairs
Quotes on the cited page
100% 2,204 of 2,204 rows, re-checked independently
Rows that are real findings
91% Estimate from a sample of 50 checked rows
Time per set
3.4 min AI reading time for an average 105-page set. Your free preview usually arrives in 5 to 15 minutes.

The bar we set before publishing: find at least 90% of findings with at least 95% of quotes correct. This run meets it.

FileAI compared with one big prompt

We also gave the same AI model each whole set in a single request with a one-line instruction: “list everything a home buyer should know before closing, with page numbers and quotes”. That is roughly what you get by pasting an HOA package into a chatbot. Both were scored the same way.

Measure FileAI Single prompt
Findings found in the association's documents 91% 193 of 212 80.2% 170 of 212
Quotes found on the cited page 100% 2,204 of 2,204 99.9% 940 of 941
Rows that are real findings estimate, 95% range 91% range 80.9–96.1% 67% range 55.6–77.2%
Rows per set 367.3 156.8
Pages per set 105.3 105.3
Minutes per set 3.4 3.2

The single prompt found about eight in ten of the findings and listed under half as many rows, and a third of the rows we sampled from it were wrong or not findings. It did well on insurance, rules and special assessments, and worst on reserve figures and repair schedules. FileAI reads each set in sections of about eight pages with written rules for every type of finding, then merges the rows and checks every quote against its page.

Where it misses

Findings found by type. The misses are listed honestly below, so you know what to double-check yourself.

Type FileAI Single prompt
Dues and fees 93.8% 15 of 16 93.8% 15 of 16
Special assessment 100% 6 of 6 100% 6 of 6
Reserves and budget 84.3% 43 of 51 66.7% 34 of 51
Litigation and disputes 100% 6 of 6 100% 6 of 6
Insurance 94.7% 18 of 19 89.5% 17 of 19
Rule or restriction 93.5% 43 of 46 93.5% 43 of 46
Major repair or project 88.5% 46 of 52 65.4% 34 of 52
Delinquency and red flag 100% 14 of 14 92.9% 13 of 14
Key date 100% 2 of 2 100% 2 of 2

Dates and ages of studies

A reserve study's own date or inspection date is sometimes not listed on its own, so a study that is six years old may not be flagged in the rows. Check the date on the first page of every study.

Single rules in long documents

A one-line ban, such as on antennas, can be left out of a long set of rules, while the larger rules (rentals, pets, parking, fines) are found. Read the rules yourself for anything you plan to do.

Totals that cover a period

A figure like a 40-year replacement total or a payment spread over several years can lose its period in the summary row. Open the page the finding links to.

Whole-building budgets and balances

Some balance sheet lines (such as deferred assessments) are summarised under another line. The headline reserve numbers, percent funded and special assessments are found reliably.

What we tested

6 sets of real public documents (23 PDF files, 632 pages) that homeowners and condominium associations publish on their own websites, in AZ, CA, DE, FL, VA, WA. They are budgets, reserve studies, governing documents, board minutes, insurance certificates and policies, from 63 to 166 pages each.

The documents are used for testing only. We don't republish them, so they are described by state and type of community, not by name.

  • AZ villa community · Minutes, insurance and reserve study

    63 pages · 27 items in the answer key

    FileAI found 88.9%, single prompt 88.9%

  • CA planned community · Budget package, fines and insurance

    89 pages · 31 items in the answer key

    FileAI found 90.3%, single prompt 93.5%

  • DE townhome community · Minutes, audit and reserve study

    78 pages · 38 items in the answer key

    FileAI found 86.8%, single prompt 68.4%

  • FL condominium · Structural integrity reserve study and declaration

    166 pages · 24 items in the answer key

    FileAI found 83.3%, single prompt 79.2%

  • VA master association · Budget, covenants and reserve study

    95 pages · 38 items in the answer key

    FileAI found 94.7%, single prompt 76.3%

  • WA high-rise condominium · Minutes, rules and reserve study

    141 pages · 54 items in the answer key

    FileAI found 96.3%, single prompt 79.6%

How we measured it

The answer key

The answer key was written by reading the documents page by page and listing what a buyer should know: 212 rows in all. It was written before FileAI was run on the documents, by an AI assistant following written rules, not by FileAI. A program then checked that every quote is on the page it cites.

No real estate professional has reviewed the key. It covers the material findings, not every line of every document, and it will have its own misses and judgement calls, so treat these numbers as a careful estimate, not a guarantee.

Most of the findings are about the association as a whole (212). Findings about one home (0) are rare because these public documents are not about a particular home; a resale certificate is.

When an item counts as found

A row has to cite the same file and a page within one page of the answer key's page, and say the same thing. A grading model decides the meaning with strict rules: a related but different item, a row that leaves out the deadline or amount, or a different number does not count.

The grader is an AI model with a strict rubric, and we did not hand-label its decisions for this kind. It is sometimes too strict: among the nineteen misses below, FileAI's rows do state at least one of them (a Florida condominium's 0.34% percent-funded figure), so the real recall is a little higher than shown.

When a quote counts as correct

The quote must appear on the page it cites, word for word, ignoring only spacing, capital letters and the style of quote marks and dashes. A quote may start on the cited page and run onto the next. Rows from tables count when every word of the quote sits together in that part of the page. This check is separate from the one FileAI runs on itself.

Extra rows and time

FileAI lists more rows than the key, because the key is selective: it holds the material items, not every line. To check those extra rows are real, we graded a random sample of 50 of them against their page.

Time is the AI processing time for a set, with sections read in parallel the way the service runs them. It leaves out waiting in the queue and sending the email, and the person who checks every row, so a breakdown takes longer than this to reach you: after you upload, the free preview is usually ready in 5 to 15 minutes, and the checked full breakdown usually within one business day.

Five of the fifty sampled rows that matched nothing in the answer key were not findings a buyer needs (for example a general procedure) or were slightly off. FileAI lists everything it can support with a quote, so expect more rows than a person would write; the preview and the ratings lead with the ones that cost money.

Limits of this test

  • 6 document sets is a small sample, and state disclosure rules differ.
  • The answer key was written by an AI assistant and not reviewed by a person, and it lists the material findings only.
  • Some public documents are scanned images with no text; FileAI can't read those yet, so they are left out of the test.
  • AI output varies a little from run to run; this is one run of each system.
  • We adjusted the prompt after reading its output on a separate public HOA document set, which is not part of this test. The six test sets were run once with the final prompt on gpt-6.1-sol.
  • Pages that are scanned images or drawings have no text we can read. FileAI lists them as pages it couldn't read, so you can check them yourself.

A person also checks every row

While volumes are low, a person reviews every row of a paid breakdown against its page before it is delivered. Every row still links to its page so you can check it yourself.