FAQ
Questions we get asked.
Honest answers, including the ones a tool would not give you.
Do our documents leave our premises?
Not during conversion. The pipeline runs on our own European servers or on your machine, and nothing is sent to any third party while documents are extracted, cleaned and chunked. If the connected assistant uses a hosted embedding model for search, that is a separate, documented choice you make, and a local model is available for customers who require no third-party processing.
Do you summarise or rewrite anything?
No. Nothing is summarised, paraphrased or filled in. Only noise is removed: page numbers, running headers, logos, decorative symbols, repeated mastheads. What the source says is what the knowledge base says. Numbers, units, currencies, dates and footnotes are preserved as written.
What happens to personal data in our documents?
It is flagged, never dropped silently. Every chunk that carries person-level data, for example a directory with named contact persons, is marked as such. Whether a public assistant may show it is your decision, taken in the decision session and written into the record with a name and a date.
Our documents contradict each other. What then?
Both statements are preserved as they stand, the conflict is listed in the report, and you decide which is current. The decision goes into the record. The connected assistant is instructed to say a fact is not listed rather than to guess or to pick one silently.
Which languages do you handle?
Greek, Dutch and English today, properly: hyphenation with a dictionary, month names in all three languages for dating documents, person-column detection in Greek and English. Search works across languages, so a Greek question finds a Dutch document.
We have scanned documents. Does that work?
Yes. Documents without a text layer are identified in the inventory and go through OCR. Every chunk that came from OCR is marked as such in the report, so you know which text was read by a machine from an image.
What exactly do we receive?
Three things: the knowledge base itself (one chunk per line with full source metadata, plus a readable version), the processing report (counts, duplicates, boilerplate, files to review, validation), and the questions only you can answer, with the record of your answers once the session is done. If we connect the knowledge base to an assistant, the delivery test results come with it.
How long does it take?
The inventory takes minutes. Conversion of a typical data set is under a week, and the curator's review is under two days of work. The decision session with you takes thirty minutes.
Can the knowledge base feed our own AI platform?
Yes. The primary output is a plain JSON-lines file with lightweight metadata, made to be embedded into the vector store of your choice. Standard connections to common vector stores and to our own assistant platform are on the roadmap; a file export always works.
How is it kept current?
New documents go into a watched folder. On a schedule, monthly or on demand, the same pipeline runs, the same report is produced, and you get a diff: new, changed and removed chunks. The test questions are run again and the answers compared.
What is a chunk?
A piece of a document, typically a section or a group of spreadsheet rows, sized so a search engine can find it and an assistant can read it. Every chunk starts with its own context (document, source, date, section) so it makes sense on its own, and carries enough metadata to trace it back to the page or the sheet it came from.
What about GDPR and the AI Act?
CODEBLACK acts as a data processor. A standard processor agreement, per-customer isolated storage, deletion on completion unless retention is agreed, and a sub-processor list are part of every engagement. Where the output feeds a customer-facing assistant, the AI Act disclosure that the user is talking to an AI is implemented and documented per delivery.
Why not just use a parsing tool?
Use one; we do. Parsing is a commodity. What a tool does not do is ask which of two versions is current, whether a contact list may be shown, or which of two member counts is right, and then take responsibility for the answer. The report, the flag lists and the decision session are the product.
Something else? Write to info@pillardelta.ai.