ai technology

Ontology: Engineering or Philosophy?

Junyoung Park · 2026-10-07 · 21 min

Ontology

Spend enough time talking about AI, and someone will tell you that you now need to study philosophy too.

AI writes the code, agents fetch the data, so what really matters is the ability to define the problem. So far, I agree. It’s what comes next that feels off. When did “defining the problem matters” become “you don’t need to understand computer science”? Now that we can give instructions in natural language, apparently writing well and having a grounding in the humanities matter more. Honestly, I’m not a fan of explaining why one field matters by pushing another out of the picture.

Ontology is an easy subject to drag into this argument. Even its name comes from the study of being. You have to think about which concepts to distinguish, but once you try to implement anything, you’re dealing with databases, queries, and logical inference. Calling it philosophy sounds reasonable. Calling it engineering sounds reasonable too. My feeling, though, is that deciding which one matters more isn’t a particularly useful exercise. We need to know what we’re trying to solve, how far our data and technology can take us, and whether we can keep the thing running. We don’t need to decide which degree wins.

That doesn’t mean I want to stop at “both are important.” When you follow what an ontology actually does, separating the two starts to look less straightforward in the first place. I’ll use a fictional manufacturer, Company A, its Product Z, and Supplier B to explain things. The companies and business figures below are all examples.

Raphael's The School of Athens, with Plato and Aristotle at the center.
Raphael, The School of Athens, 1509–1511. Photograph: Wilfredor, CC0. Image source and license

Ontology takes its name from Greek roots concerning being and study or discourse. It’s connected to philosophical questions about what exists and how we distinguish one kind of thing from another. Of course, putting Aristotle’s picture here doesn’t mean he invented RDF. In modern computer science, an ontology is closer to a model that explicitly describes the concepts and relationships needed for a particular purpose. Gruber called it an “explicit specification of a conceptualization,” which, honestly, doesn’t tell you much about what you’re supposed to do with it.

Just imagine trying to decide who counts as a customer at a company. Sales might include businesses that haven’t signed a contract yet. Finance might mean the organizations that paid. Support might mean the people actually using the product. Everyone uses the same word, but they’re referring to different things. Ask how many customers there are, and the answer will depend on which department you ask. Having a customer table in the database doesn’t fix that. They haven’t even agreed on what belongs in it.

So you distinguish prospects, contracting parties, payers, and end users, then describe their relationships. At this point, philosophical thinking seems pretty important: you’re separating concepts. But the next job is to find out which record in which system refers to each entity. You have to reconcile different IDs for the same legal entity and decide how to detect incorrect matches. Deciding what something means and handling data according to that meaning are more closely connected than they might seem. Logic and knowledge representation have worked across this boundary for a long time anyway. Ontology didn’t suddenly reverse the roles of the humanities and engineering.

So We Don’t Need CS Anymore?

Add reports about falling junior hiring and declining interest in CS degrees, and the conclusions get more dramatic. AI will do the coding, so we should study the humanities instead. There is something pretty ironic about advances in technology prompting renewed discussion of humanistic knowledge. But we should distinguish between describing actual hiring demand and talking about abilities we think are valuable.

There are real changes in the data. Stanford Digital Economy Lab’s August 2026 analysis found US employment among 22–25-year-olds in highly AI-exposed occupations roughly 19% below a comparison trajectory based on less-exposed peers. ASEE found freshman enrollment fell from 2024 to 2025 in about 91% of 135 US CS programs reporting continuously during 2018–2025. Neither figure means that AI got 19% of developers fired or that the whole world is abandoning CS. The employment analysis doesn’t establish AI’s causal effect, and enrollment changes also sit in the context of the preceding pandemic-era surge.

Jumping from there to “therefore, companies want more philosophy graduates” is a separate argument. Declining junior employment and rising humanities hiring aren’t the same statistic. I can agree that humanistic thinking matters without pretending those figures prove that conclusion. This is the part that bothers me: a remarkably complicated computational technology becomes widely used, and somehow that gets interpreted as a reason not to understand the computation.

Since AI accepts natural language, it can look as though writing good sentences is enough to make it work well. Clear instructions and appropriate examples do help. But an LLM isn’t a system that fulfills the meaning of your sentences as if they were contractual obligations. A simplified expression for an autoregressive language model looks like this:

Pθ(y∣x)=∏t=1TPθ(yt∣x,y<t)P_\theta(y\mid x)=\prod_{t=1}^{T}P_\theta(y_t\mid x,y_{<t})

Here, xx is the input—questions, instructions, reference material—yty_t is the token to generate, and θ\theta represents the learned parameters. The model continues its output by computing next-token probabilities conditioned on the input and preceding tokens. I’m not saying this equation explains everything the model can do. But we do need to understand that the ways we influence its output are different. Changing a prompt changes the input conditions. Retrieving documents through RAG supplies information to refer to. Training or fine-tuning changes parameters, while decoding settings affect how tokens are selected from the output distribution. They may all fall under “using AI well,” but they affect different parts of the system.

If an agent answers from an old contract, for example, the next step isn’t to make the instruction more polite. It’s to find out why the old document was retrieved. Lowering the temperature won’t produce an up-to-date contract. Requiring JSON won’t make the supplier ID inside it correct. When a result is wrong, you need to distinguish an ambiguous question, faulty retrieval, a bad data mapping, and a model misreading its evidence. Treat all of those as prompt problems, and you’ll just keep rewriting instructions.

That’s why I think understanding the technology itself is part of developing useful insight. Of course, not everyone needs to implement a Transformer. The depth you need depends on your role. But someone on the team has to understand the difference between probabilistic output and verified facts, measure failures, and control actions through code and permissions outside the model. A grounding in the humanities is also more about questioning assumptions, distinguishing concepts, and understanding other perspectives than writing pretty sentences. I don’t think this is a problem where learning one side lets you ignore the other.

Is Ontology a New Database for AI?

So does adding an ontology solve these problems? First, ontology isn’t a new kind of database invented when agents came along. It has a much longer history as a way to represent and share knowledge. You can build it over an RDB, but an RDB isn’t a prerequisite. It’s easier to think of it as a layer describing what the data in existing systems means and how it should be connected.

Suppose A’s ERP lists B as a VENDOR, while the production management system lists it as a PARTNER. A contract only uses its legal name. A person might recognize the company from context, but a system needs that connection made explicit. An ontology describes what a supplier is and how it relates to products and parts. It doesn’t automatically turn records with similar names into the same company. Whether they actually refer to the same entity still needs to be checked against identifiers and evidence.

It’s also a little different from a knowledge graph. A knowledge graph connects individual statements such as “Product Z uses Battery B-21” and “B supplies B-21.” An ontology defines the concepts, relationships, and premises for inference used in those statements. They often appear together, which makes the distinction confusing, but connecting specific entities and defining how to interpret those connections are distinguishable jobs.

Knowledge graph facts contrasted with ontology concepts, relationships, and rules

Drawing more edges doesn’t mean you’ve clarified the meaning. A well-designed relational schema already contains domain knowledge. That’s why I’m wary of saying a database has become smarter simply because it’s now represented as a graph. The important part is less about the shape of the storage and more about why you defined those relationships and what you can actually do with them.

I can see why agents have brought renewed attention to ontology. Once generative AI starts querying databases or registering business requests instead of just producing answers, it needs to handle the same entities and rules consistently across systems. What people call agentic AI is closer to a system combining a generative model with planning, tool use, and state management than an entirely different species of model. Ontology can help with that. It isn’t a mandatory component of every agent.

Imagine retrieving a relevant contract through RAG. Finding the document is useful, but you still need to establish whether its Supplier B is the same legal entity as B in the ERP, whether the contract is valid, and which part specification an alternative supplier’s certification covers. There is still work between fetching a document and producing an answer that fits the business question.

A comic in which an agent confidently infers a cause after reading a meeting note

Finding related information doesn’t prove a cause. The same goes for linking delivery delays and falling production in a graph. A connection and a causal relationship are different things, but it’s easy to overlook that distinction when the answer reads naturally.

How the Concepts Are Actually Represented

The same distinction is a useful starting point for TBox and ABox. A TBox contains general definitions such as “batteries are parts” and “suppliers are organizations.” An ABox contains assertions about individuals: “B-21 is a battery” and “B supplies B-21.” An RBox describes characteristics of relationships, such as whether supplying and being supplied by are inverses, or whether following a particular relationship through multiple steps preserves that relationship. Some accounts include the RBox within the TBox. This doesn’t mean you need to build three separate databases.

TBox concepts, ABox individual assertions, and RBox relationship properties

RDF is a basic model for representing this data. Its three-part statements—subject, predicate, and object—are called triples. Think “Supplier B / supplies / B-21.” RDF itself doesn’t mean an XML file format. It can also be expressed in formats such as Turtle or JSON-LD. Here’s an example in Turtle:

@prefix ex: <https://example.com/supply/> .

ex:ProductZ a ex:Product ;
    ex:usesPart ex:BatteryB21 .
ex:BatteryB21 a ex:Battery .
ex:SupplierB a ex:Supplier ;
    ex:supplies ex:BatteryB21 .

Product Z belongs to Product and uses B-21 as a part. B-21 is a Battery; B is a Supplier and supplies B-21. It’s just the example from earlier written in a specific syntax. The a abbreviates rdf:type, indicating which class something belongs to. An IRI such as ex:SupplierB identifies an entity. Having IDs doesn’t resolve duplicate records for the same company, though, so assigning and connecting those identifiers still needs to be managed.

RDFS adds vocabulary for classes and hierarchies, while OWL allows richer logical definitions, including equivalence, disjointness, and relationship characteristics. This is where the word axiom comes up. It sounds rather grand, but it’s a formal statement the model adopts as a premise. We’ve chosen to define something that way; we haven’t been given a guarantee that reality agrees. Assertions in the ABox can also be wrong if the source data is wrong.

For example, suppose this service defines batteries as critical parts, and a critical-part supplier as a supplier supplying at least one critical part. B can then be classified as a critical-part supplier using the data above. That follows from the premises. It’s a different kind of inference from an LLM generating a plausible explanation from context. But it doesn’t also establish that B is a risky supplier. Supplying a critical part and posing a delivery risk are conditions that need to be defined separately.

The distinction becomes clearer when we try to validate data. OWL follows the open-world assumption. Not having something recorded doesn’t mean we can conclude it doesn’t exist. If a supplier’s certification isn’t in the data, perhaps it isn’t certified, or perhaps that information hasn’t been connected yet. Writing an axiom that every employee has at least one manager doesn’t necessarily require an explicit manager ID in the current data either.

A real workflow, however, might need to block a review until a certification expiration date has been entered. SHACL can validate a selected RDF graph against shapes describing required values, cardinalities, and datatypes. That checks the supplied data; it doesn’t declare that we know everything about reality. “We don’t know whether it has certification” and “so we shouldn’t approve this yet” can both be true. Treat inference and workflow validation as the same thing, and the implementation will be wrong from the start.

RDFS domain and range also shouldn’t be read as ordinary database input constraints. If the domain of supplies is Supplier, a statement that X supplies something can allow X to be inferred as a Supplier. You might expect the system to reject a bad value, but it assigns a type instead. This is the kind of detail I mean when I say that knowing what the words mean is different from knowing how the system actually behaves.

Where Does the Data Come From?

Once we’ve defined the concepts, we need to retrieve actual data. SPARQL is the language used to query RDF graphs. To find critical parts used by Product Z and their suppliers, we can describe the relationship pattern like this:

PREFIX ex: <https://example.com/supply/>
SELECT ?part ?supplier
WHERE {
  ex:ProductZ ex:usesPart ?part .
  ?part a ex:CriticalPart .
  ?supplier ex:supplies ?part .
}

Notice that the earlier Turtle example only asserts the Battery type. Finding it as a CriticalPart requires storing the inferred type in advance or configuring suitable reasoning. Using SPARQL doesn’t automatically run OWL inference. Writing down definitions and building an environment that can use them in queries are separate tasks.

That doesn’t mean we have to throw away the existing relational databases and move everything into a graph database. We can copy selected data into an RDF store or access the original database through mappings. Access through an ontology is called OBDA, or Ontology-Based Data Access. R2RML describes mappings from relational data to RDF, while a supporting engine handles the actual query rewriting. Not every arbitrary OWL definition can be turned into efficient SQL. We need to consider what we can express and what it costs to execute.

Turning a natural-language question into SQL is Text-to-SQL; turning it into SPARQL is Text-to-SPARQL. We might leave large delivery-history aggregations to an RDB, find product–supplier relationships in a graph, and retrieve contract terms from documents. Those results and their sources become context for the model. RAG doesn’t have to mean vector similarity search alone.

A hybrid architecture connecting natural-language requests to graph, relational, and document retrieval

The diagram is only one possible arrangement. Independent lookups can run in parallel, but if a contract lookup requires a supplier ID first, there’s an order to follow. If an OBDA engine rewrites SPARQL into SQL, there’s no need for an LLM to generate SQL again. Adding every familiar technology to the diagram without understanding how they fit together won’t make a better system.

And a query being syntactically valid is different from its answer being correct. It may confuse dispatch dates with arrival dates or duplicate a supplier in a join. It may retrieve a contract the user isn’t authorized to see. Schema mapping, execution planning, permission checks, result validation, and provenance still have to be implemented. If the action modifies data rather than just reading it, approval and recovery from failure matter too. Being able to ask in natural language means the user sees less of this work. It doesn’t mean the work has disappeared.

First, Decide What You’re Building

This is why drawing out tacit domain knowledge is an important part of ontology design. There are things business specialists take for granted that aren’t explicitly recorded in data or documents. They’re the parts introduced with “that’s just how we do it.” Once you ask how to put them into a system, there tend to be more exceptions than you expected.

Take late delivery. Purchasing may check whether the contractual delivery date was missed. Logistics may look at the warehouse arrival date. Production may say there was no production problem because safety stock covered the delay. Compliance with the contract and risk to production are different judgments. Grouping them under a single word like “delay” or “risk” makes it hard to get the intended answer, however well the model is built.

This is where a CQ, or Competency Question, comes in. We write down what questions the ontology should be able to help answer. “Show me risky suppliers” hasn’t even defined risk yet. We can make it more specific by asking whether there were deliveries at least two days past the contractual deadline in the last 30 days, whether fewer than five days of the relevant part remain in stock, and which products are affected. If we’re looking for alternative suppliers, we’ll also need valid certifications and the sources supporting the answer.

Writing the question makes the necessary connections between delivery, inventory, products, suppliers, and certification more apparent. But we still have to decide whether five days of stock is calculated from average consumption or a confirmed production schedule. A good CQ doesn’t finish all the definitions. Linking it to actual data and expected answers can turn it into a query test; comparing it with the time and error costs of the existing process can help evaluate the benefit. The CQ itself doesn’t guarantee ROI.

An ORSD, or Ontology Requirements Specification Document, records agreements about purpose, scope, users, intended uses, and requirements. It’s an artifact proposed in the NeOn methodology, not a mandatory international standard form for every project. What matters more than the name is deciding what we will and won’t do this time. For A, that could mean starting with one domestic plant, two products, and three critical part categories. Connect ERP delivery and inventory records with certification documents, but leave out currency and weather forecasting and don’t automatically penalize suppliers. Include everything that seems relevant, and estimating the cost gets even harder.

There isn’t just one direction for design, either. Top-down starts with broader concepts. Bottom-up develops concepts from existing data. Middle-out starts with concepts needed for key CQs and expands in both directions. Whichever approach we take, we still need to go back and forth between data and questions. Well-organized concepts without the necessary data are a problem. So is a graph of existing ERP columns that can’t answer the questions we actually need answered.

Who Takes Responsibility After It’s Built?

My skepticism about ontology is less about whether we can build a model and more about how hard it is to estimate what building and maintaining it will take. If meanings diverge across systems, relationships have to be followed repeatedly, and multiple services reuse the same rules, I can see the value. But a fixed sales report may be fine with a SQL view, and questions about policy documents may be solved by improving search. Being able to do something with ontology isn’t a reason we have to do it with ontology.

Initially, there are interviews, agreement on terminology, and identifier cleanup. Mappings, permissions, and tests need to be built too. Once it’s running, source databases change, data errors come in, and rules change. The cost includes not only reasoning, storage, and model usage but also the time people spend understanding and implementing those changes. Counting classes and checking a graph database’s price doesn’t tell us the total effort involved.

Suppose A changes its delivery rule from dispatch date to warehouse arrival date, excludes weekends, and makes exceptions for emergency orders. It looks like a definition change, but first we have to find out whether actual arrival timestamps exist in the data. We need a holiday calendar and changes to mappings, calculations, and tests. We also have to decide whether to recalculate historical reports using the new rule. Compare numbers under the same label of “delay” when the definition has changed, and a policy change could be mistaken for improved performance. That’s why versions, effective dates, and sources need to be recorded.

If the necessary data doesn’t exist, we’re back to talking with the business team. Do we use a proxy, change the input process, or narrow the service’s scope? It isn’t a matter of completing the philosophical definition and handing it to an engineer. Defining, implementing, and checking keep feeding back into each other.

A comic about organizational, schema, and permission changes arriving after a successful demo

How useful is it to ask which field matters more in that situation? We definitely need to know who approves changes to concepts, who updates mappings, and who handles deployment and rollback. Separating responsibilities matters. But that doesn’t need to become an argument about which discipline is the substance and which is just an accessory. We can identify what’s blocking us right now without deciding which field will always matter more.

Of course, I’m not saying one person has to be good at all of it. Business specialists, modelers, data engineers, and ML and service engineers need to communicate the assumptions behind their judgments and the limits of what they can do. Knowing where your knowledge stops is a different requirement from being an expert in everything.

That’s why I don’t think “we have to build it and use it to find out” is inherently wrong. Some problems only appear when you try. But it shouldn’t become an answer that dodges questions about costs and benefits. If we’re uncertain, we should test on a small scale. For A, I’d start with around 15 CQs and compare the existing process, a SQL or RAG alternative, and an ontology-based approach using the same data snapshot and permissions. Alongside accuracy, I’d check whether the evidence is correct, how long people spend rechecking results, response times, and the hours needed to update rules and mappings. Those numbers describe a proposed experiment, not actual results.

It also matters that we don’t credit ontology with every benefit of data cleanup. If deduplicating supplier codes improved the result, we should test the cleaned data in the existing approach too. Where possible, compare the same data with and without the ontology layer. And if it isn’t better than the alternatives on essential questions, or maintaining it takes more time than it saves, we should be willing to stop expanding it. All the more so if nobody has agreed to own the rules. Being able to build something and having a reason to operate it are different things.

One Last Thing

What I came back to while working through ontology was the fairly obvious idea that I need to understand the problem I’m solving. But understanding the problem doesn’t just mean writing the requirements clearly in natural language. We need to decide what counts as the same entity, know what data can actually be observed, and understand what we can compute from it. Then, when the result is wrong, we need to be able to tell where it went wrong.

I’m not saying humanistic thinking doesn’t matter. It matters enough that I’d rather not see it reduced to “being good at talking to AI.” On the other hand, knowing engineering doesn’t automatically tell us what people and organizations want. A mathematically sound model can still be useless if it was solving the wrong problem to begin with. This isn’t a setup where mastering one side lets us replace the other.

That’s why I think it’s premature to look at difficult junior hiring and conclude there’s no reason to study CS. The easier it becomes to produce results, the more we need knowledge that lets us judge them. Juniors need to define real problems, implement things, and experience failure to learn that judgment too. “AI does it, so I don’t need to know” may get you an output. It won’t tell you why the output is wrong.

Oh, so is ontology philosophy or engineering? I don’t really see why we have to pick one. We need to understand what the work requires, what the technology can support, and how to change the system when either of those changes. Instead of deciding which field matters more, wouldn’t it be more useful to ask whether I’m capable of doing that work?