Summary
This article comes from a real legacy system: more than two million lines of C and C++, a large PHP and PL/SQL layer, and hundreds of documents spread across releases that contradict each other. I had spent weeks tuning retrieval pipelines on it. Vector search kept returning chunks that looked alike, taken from different releases of the same specification, and the model merged them into answers that looked right and were not. The turning point was accepting that the problem was not retrieval but representation.
Ontology-based context engineering means building a machine-readable map of the domain before giving anything to the model, and using that map to decide what the model is allowed to see. Embeddings capture similarity; an ontology captures meaning. Facts from the graph select the artifacts, the documents and the release; similarity search then finds the right passage inside that selection. Both techniques are used, and the whole design lies in that order.
The ontology I ended up with is small: about a dozen classes and ten relations, from C modules and PL/SQL packages to specification documents and business rules, with a named graph per release so that a reasoner never mixes facts from two versions. A few hundred lines of Turtle were written by hand; deterministic collectors generated the rest, a SHACL gate quarantined what did not validate, and every change goes through a pull request, like code.
RAG is not dead, it is just not sufficient when the meaning of a system is split across code, database, documents and releases. RDF, RDFS, SKOS and SHACL turn out to be a pragmatic stack, with OWL kept for the places where inference pays. The upfront modelling cost real time, far less than the time I had lost on pipelines that could not work, and every answer now comes with a path a developer can verify.
Key ideas
- Embeddings capture similarity, ontologies capture meaning: the graph decides which artifacts and which release are on the table, and similarity search finds the passage.
- A legacy system’s meaning is never in one place; a question about an invoice spans a C module, PL/SQL packages, a table and two contradicting specifications.
- Keep the ontology small and treat it as code: a dozen classes, ten relations, Turtle files in Git, SHACL validation and pull requests.
- One named graph per release, so that a reasoner cannot mix the facts of 2022 with those of 2024.
- Start with one subsystem: it proves the value in weeks, and the ontology grows from there.
Why I wrote this
It follows my articles on context engineering and on BMAD: I wanted to take the subject one step further, from the workflow to the representation. I did not have the luck of a modern, well-documented application; on this system, the approach started the moment I accepted that the problem was representation rather than retrieval.