DATA INTELLIGENCE

You do not have data.You have documents. CODEBLACK turns a company's document collection into a clean, verified, searchable knowledge base, and keeps it that way.

Why CODEBLACK

Your documents know more than your systems do.

Years of contracts, reports, newsletters and spreadsheets. Several languages. Duplicates, old versions, names that must stay inside. Point an AI at that pile and it answers from the wrong file.

CODEBLACK turns the pile into one verified knowledge base. What is current, what is private, what is true: sorted by us, confirmed by you in thirty minutes, traceable to the page. Then it feeds whatever you build on it. An assistant. A search. An audit. A data room.

Nothing is made up.

Codeblacked
250 MB9 MB

250 MB of a trade association's archive became a 9 MB knowledge base. Its assistant answers from it today, in three languages, with sources. Their next system starts from the same base.

How it works

Watch a folder become a knowledge base.

May a public assistant show the contact persons?
keepdrop
Two newsletters are 91% the same. Which one counts?
March 2024April 2024
The FAQ says 130 members, the history says 134.
130134
FAQ is current; history marked out of date
recorded
14 May 2024, in The Hague. Registration closes on 7 May.
Newsletter April 2024 · p. 1
450 EUR per year.
Statute · Article 3
That is not listed in the documents.
nothing invented
The problem

Every AI project stalls in week one, between the vendor and the folder.

A company that wants a chatbot, an internal search or any assistant discovers the same thing: it has folders of PDFs, Word files and spreadsheets, accumulated over years, in mixed languages, full of logos, page headers, duplicate versions and half-empty tables. AI tools cannot use that.

The vendor assumes the data is clean

Parsing libraries and cloud document services are commodities. They extract. They do not ask which version is current or whether a contact list may be shown to the public.

The company assumes the vendor will handle it

Nobody inside has time to read 150 files with a checklist, and outside tools do not raise the questions that need an answer.

The hard part is judgement, not extraction

Boilerplate or content. Current or stale. Public or internal. Two documents that disagree on the member count. A model that invents a city for an empty address field.

The offer

Four blocks. Sold separately or together.

Convert

Every document, extracted and cleaned

PDF, Word, Excel and legacy Office files extracted, cleaned without rewording, structured, chunked and deduplicated into a machine-readable knowledge base with full source metadata. Plus a readable version, a source inventory and a processing report.

Curate

A human review of everything flagged

Duplicates, conflicting facts, personal data, stale versions, unreadable files, sensitive strings. Delivered as a short list of decisions for you, and a written record of what was decided, by whom, on which date.

Connect

Wired into where it is used

A customer-facing assistant, an internal search, or your own AI platform, with retrieval that cites sources and refuses to invent.

Keep current

It never drifts from the documents

New documents dropped into a folder are processed, reviewed and re-indexed on a schedule, with a report of what changed.

You see three deliverables: the knowledge base, the report, and the questions only you can answer.

How it works

Seven steps. Under a week for a typical data set.

  1. Intake. A short questionnaire and a folder: languages, audience, what the assistant must answer, who signs off on personal data and conflicts.
  2. Inventory. Every file counted, typed and sized. Scanned documents, images and person-level columns identified before anything else runs. This is what the quote is based on.
  3. Conversion. Per-format extraction, cleaning without rewording, structure recovery, token-aware chunking with context on every chunk, deduplication at document, paragraph and chunk level, validation.
  4. Review. A curator works through the report's flag lists with a checklist built from experience, samples chunks against the originals, and prepares the decision list.
  5. Your decisions. A thirty-minute session: personal data yes or no, which version wins, which documents are out of date, what stays internal.
  6. Delivery. Knowledge base and report handed over, or connected to an assistant and verified with test questions whose answers only the documents contain.
  7. Refresh. Monthly or on demand: the same pipeline, the same report, a diff of what changed.

Processing runs on our own European infrastructure or on your machine. Nothing is sent to third parties during conversion. If retrieval uses a hosted embedding model, that is a documented, separate choice you make.

CODEBLACK on site

Comes to your server. Does the job. Leaves. Your documents stay.

For archives that must not leave the house. CODEBLACK runs as a sealed container on your own server, with the network switched off. It writes the knowledge base, the report and the flags to a folder you own, and removes itself.

Sealed

No network inside the container, enforced twice: by the container and by the runner itself. An attestation lists the versions, the hashes and every file written, so the claim can be checked rather than believed.

Only the report travels

The curator needs the report and the flags: file names, counts and short excerpts, a few hundred kilobytes. Your documents and your knowledge base stay where they are.

Your decisions, as always

What is current, what is private, what is true: answered by you in the portal, applied on site, recorded with a name and a date.

What makes it different

The judgement layer is the product.

Judgement calls become recorded decisions

Personal data, conflicting facts, which version is current: each is a question put to you, answered with a name and a date, and applied from that record.

Verified, not assumed

Every delivery ends with a test: questions whose answers exist only in the documents, asked to the connected assistant, with the answers and their sources logged.

Nothing invented

No summarising, no rewording, no filled gaps. What the source says is what the knowledge base says, and the assistant is instructed to say “not listed” rather than guess.

European and multilingual

Greek, Dutch and English handled properly today, including cross-language search: a Greek question finds a Dutch document. Processing in the EU, processor agreement standard.

Proof, not a demo

A live assistant answering from 2,005 chunks with sources, in three languages, is the reference.

Proof, not a demo

One trade association. 151 files. 2,005 chunks. Three languages.

A trade association handed us its whole archive: newsletters, 25 magazine issues, market studies, member directories, statutes. Extraction took five minutes. Getting it right took two days of decisions.

Source151 files, 246 MB: 97 PDFs, 41 spreadsheets, 10 Word documents, 3 images
Output2,005 chunks, 8.25 MB, 96.6 percent smaller
Processingabout 5 minutes of machine time, fully local
Flags raised3 unusable files, 3 duplicate documents, 15 boilerplate blocks, 701 chunks with personal data, 1 factual conflict between documents, 25 magazine issues for reading-order review
Connected assistantanswers in Greek, English and Dutch with source and date; retrieval adds about 0.2 s in text and under 0.5 s in voice

Magazine layouts came out one word per line

Fixed with a rule for justified text, then checked by sampling against the originals.

The FAQ and the history disagreed on the member count

Both preserved. The association decided which is current. The record says so.

Contact persons in 30 directories

Flagged per chunk, never dropped silently. The association's sign-off is on file.

Who it is for

100 to 5,000 documents. 50 MB to 5 GB.

Below that a company can do it by hand. Above that it is a different kind of project.

  • Associations, chambers and professional bodies with newsletters, magazines, member directories and statutes.
  • Small and mid-size companies that want a customer assistant or internal search over manuals, price lists, policies and reports.
  • Trade offices, embassies and agencies that publish market studies and answer the same questions repeatedly.
  • Professional firms building internal knowledge search, where accuracy and traceability matter more than volume.
  • AI vendors and agencies that need clean input for their own products and would rather subcontract the preparation.

Nothing invented. Everything traceable.

Send us a folder. The inventory takes minutes, costs nothing, and tells both of us what the work is.