An AI startup · Data ontology
AI Systems for Institutions That Run on Records.
We write the ontology, send the agents that fill it, and build the system your staff work in. Then we run it.
The stack
Four Layers, All of Them Ours
This is normally four separate suppliers, plus a systems integrator hired to connect them. We do all four parts ourselves.
-
L1
Ontology
We write down what your data means
Our software reads the systems you already run and writes the first draft: what each record is, how it links to the others, and who has to sign off. We set the rules it follows and check what it produces. You end up with one document your staff can read and your software can execute.
Reads fromSAPSQL ServerTallyScanned PDFsMongoDBOracleSharePointMS AccessLegacy portalsREST APIsPostgreSQLExcelCSV exportsPaper registersGoogle Sheets -
L2
Agents & engineers
We sit with your experts, in software or in person
Much of what decides an outcome was never written down anywhere. It is in the head of the officer who has done the job for eleven years. Getting it out of there is the same job either way: ask how the work is actually done, and write the answers into the same structure as everything else. When that officer is transferred, the written version stays behind.
Sits withRevenue officersDraftsmenField engineersCuratorsLibrariansArchivistsEditorsDesk officersCase workersCompliance leadsProcurement staffExaminersRegistrarsSection officersSurveyorsForward deployed agents Software, alongside your staffIt works through the systems your people already use, asks its questions there, and writes what it learns straight into the ontology.
Forward deployed engineers Our own people, in your buildingThey sit in the room with your officers for as long as it takes, and write the same structure by hand where software cannot reach.
Which of the two we send depends on the engagement, and on some it is both. Tell us the problem and we will say what it would take.
-
L3
Application
We build the system your staff work in
Either we build it, or we put the layer inside the software you already use. Either way your staff get an answer they can act on, with the source record attached to it.
Ships asA single-window portalAn officer dashboardA search interfaceA reporting packA case file systemAn open APIA review workflowAn alerting rule setAn editorial queueA layer inside your ERPA public platformA verified roster -
L4
Operation
We run it, and we keep running it
Five government portals, a hundred-year-old archive, a publishing line and a newsroom engine are live on this stack right now, and we operate them. We do not hand over the code and leave.
Live today5 government portals10M words of research506 writers resolved25 centuries covered115,652 records classified5 publishing lines1.3M passages addressable36,472 entities2,000+ works catalogued1 newsroom engine363,042 quotations13,444 typed citations
The discipline
What an Ontology Is
Ask two departments what a pending case is and you will get two answers. Both are right, because each was defined for a different purpose years ago, and neither definition was ever written down. So the two produce different numbers and nobody can say which one is correct.
An ontology closes it. It is a plain document that says what each thing is, which one is the real number, and who decides. Once it exists, your staff and your software are reading the same definitions, and every answer can show the record it came from.
Data sources
What you have
- Scanned PDFs
- Spreadsheets and CSV files
- Databases
- Old systems still in use
- SAP and other enterprise software
Wherever your data sits, we write the connection to it. Nothing is moved or re-typed.
Logic sources
What you know
- Formulas buried in spreadsheets
- Code you already run
- Machine learning models
- Written rules and prompts
Much of it is in no system at all. Forward deployed agents and engineers sit with your experts, ask how the job is really done, and write it down in the same structure as everything else.
Systems of action
What you do
- Your ERP
- Your existing portals
- A new system we build
- An AI agent
We build the system you act in, or put the layer inside the one you already use. The result is a step someone can actually carry out.
Those three parts together are the ontology. It is kept in one place, in one structure, and both your staff and your software read it.
Why it repeats
One Layer, Rebuilt Against Every System
Four unrelated industries. No two of them stored their data the same way. We rebuilt the same layer against each one.
Case studies
Five Systems in Production
Five industries, one method: write down what the data means before building anything on it. Clients are described rather than named.
Five state portals, each on different technology and each storing its data differently. One layer answers an official’s question from that portal’s own records, and can write nothing back.
Read the case study → 115,652records classified on one portalIt only reads. It cannot alter a record even if it is asked to. Think tank A Century of Writing, Searchable by Who Said WhatTwo thousand scanned PDFs became a library you can ask who wrote about whom, on what subject, in what words. Names are matched across scripts and spellings, so the same person is not counted as several people.
Read the case study → 1,000+historical works cataloguedEvery claim is pinned to a quotation the software can find again, or it is thrown away. Publishing Graphic Stories at Scale, Every Line Traced to a SourceResearch turned into a corpus a writer can navigate, then into finished books: script, characters, backgrounds, pages a printer can use. Five publishing lines run on it.
Read the case study → 10Mwords in the research corpusA gate refuses the script if any citation does not resolve to a real line. Media A Newsroom That Runs Unattended and Publishes NothingThe engine reads live signals, plans its own slate, writes in the publisher’s voice and checks itself against sources it fetches. Then it stops, and an editor decides.
Read the case study → 0credentials that can publishOn its worst day it fills a queue rather than embarrassing the masthead. Education Handwritten Registers Turned Into Verified RowsRegistration sheets read at scale and reconciled against the paper’s own arithmetic. A model reads the page and code does every sum, so the same reading is never used to check itself.
Read the case study → 2independent readings of each totalAnything it cannot verify comes back marked unverified rather than marked correct.Use cases
Built to Order, by Institution
Grouped by the institution asking. These are capabilities, not finished projects.
Manuscripts and oral history captured properly, a public platform over the collection, an open interface any AI can read under your rules, and a picture of who actually cites you.
See the use cases → Government A different use case per departmentIndustry, revenue, procurement, drafting, disaster management, local bodies, constituency offices and public sector oversight. Each has its own problem, and the page says which.
See the use cases → Media and publishing From research to a story ready to runCoverage that stops at an editor, a research desk that turns documents into angles, analytics that answer editorial questions, and discourse measured claim by claim.
See the use cases → Legal practices The office keeps its judgementJudgments translated to filing standard under a glossary that is binding, case files structured into positions and authorities, and conflicting authority surfaced across a body of law.
See the use cases → Defence and security Runs where nothing may leave the buildingDoctrine and standing orders made answerable, assessment that keeps source reliability separate from what was reported, and every part able to run fully disconnected.
See the use cases → AI companies and labs Messy material turned into training-grade structureDomain corpora typed and given provenance with the licence position recorded, plus evaluation sets a model cannot already have memorised.
See the use cases →Open source
The Method, in Public
We publish the method where anyone can inspect it.
Our open library of the world's philosophical, classical and religious texts: more than 2,000 works split into more than 1.3 million passages that can each be quoted and cited exactly.
Visit falsafa.ai →A map of the collection drawn from the texts themselves: the people, ideas, places, groups and events across twenty-five centuries, and every time one text cites another, including whether the later author agreed or argued back.
The software that builds the Atlas is not allowed to write quotations. It can only point at paragraphs, and the words are attached afterwards from the source text.
Thothica works with some of India's largest publishers, with state governments, and with leading think tanks and small businesses.

