DATA INTELLIGENCE
You do not have data.You have documents. CODEBLACK turns a company's document collection into a clean, verified, searchable knowledge base, and keeps it that way.
Your documents know more than your systems do.
Years of contracts, reports, newsletters and spreadsheets. Several languages. Duplicates, old versions, names that must stay inside. Point an AI at that pile and it answers from the wrong file.
CODEBLACK turns the pile into one verified knowledge base. What is current, what is private, what is true: sorted by us, confirmed by you in thirty minutes, traceable to the page. Then it feeds whatever you build on it. An assistant. A search. An audit. A data room.
Nothing is made up.
250 MB of a trade association's archive became a 9 MB knowledge base. Its assistant answers from it today, in three languages, with sources. Their next system starts from the same base.
Watch a folder become a knowledge base.
Every AI project stalls in week one, between the vendor and the folder.
A company that wants a chatbot, an internal search or any assistant discovers the same thing: it has folders of PDFs, Word files and spreadsheets, accumulated over years, in mixed languages, full of logos, page headers, duplicate versions and half-empty tables. AI tools cannot use that.
The vendor assumes the data is clean
Parsing libraries and cloud document services are commodities. They extract. They do not ask which version is current or whether a contact list may be shown to the public.
The company assumes the vendor will handle it
Nobody inside has time to read 150 files with a checklist, and outside tools do not raise the questions that need an answer.
The hard part is judgement, not extraction
Boilerplate or content. Current or stale. Public or internal. Two documents that disagree on the member count. A model that invents a city for an empty address field.
Four blocks. Sold separately or together.
Every document, extracted and cleaned
PDF, Word, Excel and legacy Office files extracted, cleaned without rewording, structured, chunked and deduplicated into a machine-readable knowledge base with full source metadata. Plus a readable version, a source inventory and a processing report.
A human review of everything flagged
Duplicates, conflicting facts, personal data, stale versions, unreadable files, sensitive strings. Delivered as a short list of decisions for you, and a written record of what was decided, by whom, on which date.
Wired into where it is used
A customer-facing assistant, an internal search, or your own AI platform, with retrieval that cites sources and refuses to invent.
It never drifts from the documents
New documents dropped into a folder are processed, reviewed and re-indexed on a schedule, with a report of what changed.
You see three deliverables: the knowledge base, the report, and the questions only you can answer.
Seven steps. Under a week for a typical data set.
- Intake. A short questionnaire and a folder: languages, audience, what the assistant must answer, who signs off on personal data and conflicts.
- Inventory. Every file counted, typed and sized. Scanned documents, images and person-level columns identified before anything else runs. This is what the quote is based on.
- Conversion. Per-format extraction, cleaning without rewording, structure recovery, token-aware chunking with context on every chunk, deduplication at document, paragraph and chunk level, validation.
- Review. A curator works through the report's flag lists with a checklist built from experience, samples chunks against the originals, and prepares the decision list.
- Your decisions. A thirty-minute session: personal data yes or no, which version wins, which documents are out of date, what stays internal.
- Delivery. Knowledge base and report handed over, or connected to an assistant and verified with test questions whose answers only the documents contain.
- Refresh. Monthly or on demand: the same pipeline, the same report, a diff of what changed.
Processing runs on our own European infrastructure or on your machine. Nothing is sent to third parties during conversion. If retrieval uses a hosted embedding model, that is a documented, separate choice you make.
Comes to your server. Does the job. Leaves. Your documents stay.
For archives that must not leave the house. CODEBLACK runs as a sealed container on your own server, with the network switched off. It writes the knowledge base, the report and the flags to a folder you own, and removes itself.
Sealed
No network inside the container, enforced twice: by the container and by the runner itself. An attestation lists the versions, the hashes and every file written, so the claim can be checked rather than believed.
Only the report travels
The curator needs the report and the flags: file names, counts and short excerpts, a few hundred kilobytes. Your documents and your knowledge base stay where they are.
Your decisions, as always
What is current, what is private, what is true: answered by you in the portal, applied on site, recorded with a name and a date.
The judgement layer is the product.
Judgement calls become recorded decisions
Personal data, conflicting facts, which version is current: each is a question put to you, answered with a name and a date, and applied from that record.
Verified, not assumed
Every delivery ends with a test: questions whose answers exist only in the documents, asked to the connected assistant, with the answers and their sources logged.
Nothing invented
No summarising, no rewording, no filled gaps. What the source says is what the knowledge base says, and the assistant is instructed to say “not listed” rather than guess.
European and multilingual
Greek, Dutch and English handled properly today, including cross-language search: a Greek question finds a Dutch document. Processing in the EU, processor agreement standard.
Proof, not a demo
A live assistant answering from 2,005 chunks with sources, in three languages, is the reference.
One trade association. 151 files. 2,005 chunks. Three languages.
A trade association handed us its whole archive: newsletters, 25 magazine issues, market studies, member directories, statutes. Extraction took five minutes. Getting it right took two days of decisions.
| Source | 151 files, 246 MB: 97 PDFs, 41 spreadsheets, 10 Word documents, 3 images |
|---|---|
| Output | 2,005 chunks, 8.25 MB, 96.6 percent smaller |
| Processing | about 5 minutes of machine time, fully local |
| Flags raised | 3 unusable files, 3 duplicate documents, 15 boilerplate blocks, 701 chunks with personal data, 1 factual conflict between documents, 25 magazine issues for reading-order review |
| Connected assistant | answers in Greek, English and Dutch with source and date; retrieval adds about 0.2 s in text and under 0.5 s in voice |
Magazine layouts came out one word per line
Fixed with a rule for justified text, then checked by sampling against the originals.
The FAQ and the history disagreed on the member count
Both preserved. The association decided which is current. The record says so.
Contact persons in 30 directories
Flagged per chunk, never dropped silently. The association's sign-off is on file.
100 to 5,000 documents. 50 MB to 5 GB.
Below that a company can do it by hand. Above that it is a different kind of project.
- Associations, chambers and professional bodies with newsletters, magazines, member directories and statutes.
- Small and mid-size companies that want a customer assistant or internal search over manuals, price lists, policies and reports.
- Trade offices, embassies and agencies that publish market studies and answer the same questions repeatedly.
- Professional firms building internal knowledge search, where accuracy and traceability matter more than volume.
- AI vendors and agencies that need clean input for their own products and would rather subcontract the preparation.