Writing

From Legacy Artifacts to Served Context: My Six-Stage Method

Building a Knowledge Graph and Keeping It True, Release After Release

Summary

An ontology with no data is a diagram on a wall. This article shows how to bring it to life in six stages: analyzing, collecting, transforming, versioning, retrieving and serving. The method is applied to a legacy system mixing C, C++, PL/SQL, PHP and twenty years of bilingual documentation.

The first four stages build and maintain the graph in a CI job; the last two run at question time. Only one principle guides the whole: facts extracted by deterministic tools are considered reliable, while the proposals of an LLM remain to be validated. The LLMs run locally, with no data leaving the network.

Analyzing first defines the standards and the modeling granularity. Collecting then relies on deterministic tools: Tree-sitter for the C calls, the Oracle data dictionary for the PL/SQL dependencies and nikic/php-parser for the PHP source code. The documents are the only area where the LLM steps in: it proposes identities, versions and links, then validated by an expert.

Nothing enters the graph directly. Entity resolution groups the synonyms into bilingual SKOS concepts, SHACL validates the structure of the facts and RDF-star keeps their provenance. Versioning relies on one graph per release and facts appended then closed, with the triple store rebuilt from the repository and the fact store.

At question time, the graph selects the context before the LLM: entity linking, a SPARQL traversal filtered by release, then vector or exact search within the retained sections. Two cases — the diagnosis of a bug and a product question spanning two releases — illustrate the complete workflow, along with four objections, including the risk of a review queue that becomes unmanageable.

Key ideas

  • Deterministic facts are trusted, LLM proposals are suspected until validated: parsers and the data dictionary write to the graph, the LLM waits in a review queue.
  • The graph stops where the text begins: model the procedure, the table, the template and the document section, and keep a pointer below that floor.
  • Nothing enters the graph directly: a SHACL gate passes facts on structure, quarantines the rest with its report, and every accepted fact keeps its provenance.
  • A graph that was true at release N and never updated is worse than no graph: one named graph per release, facts closed rather than deleted, a store rebuilt, never patched.
  • The retrieval channels are not peers: the graph decides which artifacts, documents and release are on the table, and text search only reads inside that selection.

Why I wrote this

The previous article closed on the six stages as a one-page overview; this one is the method itself, each stage in the detail that overview could not give, and both runs of the pipeline, the initialization that builds the baseline and the update that keeps the graph true at every release. It stops at the gate: how the type of a question decides which channels even run is for the next article.