600 research PDFs that cannot leave the building

The volume is a problem. The confidentiality agreement is the thing that ends the conversation, and it is the reason this job exists at all.

By Cory FechnerAugust 24, 20265 min readBeyond the chat window

The short answer

Six hundred mixed-quality PDFs become one spreadsheet, a search that answers plain-English questions across all of them, and a written comparison. The files go to a machine the coordinator controls and are destroyed with it, which is what makes the job possible under her agreement at all.

A research coordinator with six hundred study documents. Some clean exports, some scans of scans, most with tables and charts that matter more than the prose around them. This is the second of four jobs in what AI can do that a chat window cannot, and it is the one where the deciding constraint is not volume.

She needs three things: the same handful of fields pulled from every study into one spreadsheet, a search that works across all six hundred at once, and a written summary of where the studies agree and where they do not.

Why the chat window stalls

The volume is a real problem — six hundred documents at a few minutes each is a fortnight of afternoons — but it is not the thing that ends the conversation.

The agreement she signed is. These documents cannot be uploaded to a third-party service, full stop. Not to a chat tool, not to a document-processing vendor, not to anything that takes custody of the file. That single clause removes every option that begins with "just upload them to".

A machine of her own is different in kind rather than in degree. The files go somewhere she controls, the work happens there, and the machine is destroyed when she is done. This is the case where the answer is not "running is faster" but "running is the only version of this that is allowed to happen".

The second problem is that the work is not one task. Reading a page, pulling a field out of it, and building a search over the whole set are three different capabilities, and no chat window does all three across six hundred files.

The models that do the work instead

Four models, chained, each handing its output to the next.

PaddleOCR-VL reads the pages, including the scanned ones, and keeps tables as tables instead of flattening them into scrambled text. That structural point is the difference between a spreadsheet where the numbers line up and one where they do not.

Qwen3-VL handles the pages OCR struggles with — charts, figures, anything where the meaning lives in the picture rather than in the words. A bar chart with no data table underneath it is invisible to a text reader and legible to a vision model.

Qwen3 pulls the specified fields out of every study into the same structure, then writes the cross-study comparison at the end. This is the language model in the chain, and it is worth noticing that it is one component rather than the whole system.

BGE-M3 builds the search index, so a question asked in plain English finds the right paragraph in the right study rather than the right filename.

The run, step by step

The run $6.40 of $25.00 cap
  1. STEP 1Read every page1 hr 5 min
  2. STEP 2Rescue the hard pages40 min
  3. STEP 3Pull the fields2 hr 20 min
  4. STEP 4Build the search index25 min
  5. STEP 5Write the comparison30 min
Machine time5h 00m
Your attention~35 min
Cost of the run$6.40

Pulling the fields is nearly half the run on its own, because it is the step that touches every page of every document and has to be careful rather than fast. Reading the pages is the next largest. Building the index and writing the comparison are quick by comparison — they work on text that is already clean.

Her part is ninety seconds: say what she needs out of them, point at the files, approve the price. The thirty-five minutes in the ledger is checking a sample of extracted rows against the source documents, which is the right place to spend attention on a job like this.

What you will have when it is done

A spreadsheet with one row per document and the fields you named. A search box that answers questions in plain English across all of them. And a written summary of the patterns across the whole set — where the studies agree, where they do not, and where the sample is too small to say.

The summary is the part that reads as magic and should be treated with the most suspicion. It is produced by a model reading six hundred extractions, and it is a starting point for a person who knows the field, not a conclusion. The spreadsheet and the search are the durable outputs; the summary is a first pass.

Why confidentiality changes the arithmetic

Everywhere else in this guide the crossover between asking and running sits around forty items. Here it sits at one.

If a document cannot be uploaded, the number of documents is irrelevant to the decision. A single confidential contract has the same answer as six hundred: the work has to happen somewhere you control, or it does not happen. Volume then only decides whether it is worth doing at all, not how.

That is worth separating out, because it is the reason this shape of work shows up in regulated industries long before it shows up anywhere else. Legal, health, HR and anything under a client NDA all hit the same wall, and they hit it on document one.

The thing that goes wrong

Extraction fails quietly. A field that could not be found comes back empty, and an empty cell in a spreadsheet of six hundred rows looks like a document that did not have that field rather than a page the reader could not parse.

So the run distinguishes between the two, and pages that defeated both readers come back flagged. Checking thirty flagged documents by hand is an afternoon. Discovering six months later that forty rows were silently blank is a different kind of problem, and it is the one that makes people distrust the whole approach.

How it works shows this same job as a handover, with your three steps set against the five that follow.

Questions people ask

Where do the documents actually go?

To the machine started for the job and nowhere else. No third-party document service sees a page, and no request leaves the session carrying file contents. The machine and its storage are destroyed when the work ends, so there is no copy of the archive left behind afterwards.

Will it read bad scans?

Mostly. Modern page-reading models handle crooked, low-contrast and photographed pages far better than older OCR, and a second model picks up the pages the first one struggles with. Pages that defeat both come back flagged rather than silently returning empty rows in your spreadsheet.

What happens to tables?

They stay tables. This is the part that separates a useful result from a useless one: older tools flatten a table into a stream of text, and the relationship between a row and its heading is lost. Keeping structure is why the numbers in your spreadsheet line up with the right study.

Can I ask questions across all of them at once?

Yes. Alongside the spreadsheet you get a search index built over the whole set, so a question asked in plain English finds the right paragraph in the right document. That is a different capability from extraction and it is built by a different model in the same run.

Is five hours accurate for 600 documents?

It is illustrative, modelled on a mixed-quality set of that size rather than measured on our own runs. Page count, scan quality and how many fields you want extracted all move it. Every session shows the price and a spend cap on screen before anything starts.

EARLY ACCESS

Renting opens shortly. Want to know when?

We are opening rentals to a small group at a time. The first 100 accounts keep the rate they join at for twelve months, whatever our prices do afterwards. Leave an email and we will tell you when your turn comes. One message, no newsletter.

Own a GPU that sits idle? Put it to work. The average provider earns $180 to $420 per month per card.

List your GPU