본문으로 건너뛰기
5 min

Graph

Explains which questions only relations can answer and how graph traversal differs from a table query.

Define entities and relations, fill them with real data, and a graph takes shape. A graph is not a picture drawn with dots and lines but a data structure you can move through by following relations.

A table query is strong at finding rows that match a condition and weak at questions that reach several steps away. For a question like "who is two hops away from Park Su-min through projects", following relations is shorter and more accurate than stacking joins.

Where table queries fall short

Tables handle questions that stay inside one table. Finding the projects that started this month needs no graph.

They weaken when the distance is not known in advance. Start from Park Su-min, move to the projects they joined, then to the employees on those projects, then to the other projects those employees joined, and a table needs one more join at every step. A question with an unknown number of steps cannot fix the number of joins beforehand. In a graph the same question becomes one act of moving along connections.

One thing stands out here. The model holds no relation linking one employee directly to another.

Loading the diagram. Mermaid source:

flowchart LR
    accTitle: Three registered relations and the unregistered link between employees
    accDescr: Employees belong to a team and join projects, and projects work with clients. Those three relations are registered in the model. No relation links one employee directly to another, so having worked together only emerges from reading two participation records that share a project.
    E1[Employee Park Su-min] -->|belongs to| T[Team]
    E1 -->|joins| P[Hanbit project]
    E2[Employee Lee Ji-ho] -->|joins| P
    P -->|works with| C[Client]
    E1 -.->|worked together · not registered| E2

The dashed line is the connection nobody registered. That two people worked together is stored as a row nowhere; it appears when two participation records sharing a project are read together. That is what a graph does.

Widening the search one step at a time

Take one object as the starting point and the connections attached directly to it form the first step. Widen the range one step further and objects two or three connections away come into view. Where to stop is decided by the question.

What separates this from a join is that the path does not have to be designed in advance. You can start without knowing which way things lead and follow connections as they appear, which is especially useful while investigating data.

The direction of a relation fixes what the connection means, not whether a neighbour is visible. A relation stating that an employee joins a project still appears as a neighbour when you start from the project; only the distinction between the joining side and the joined side stays in place.

A graph knows only what was loaded

Defining a relation in the model does not fill the graph. Relations need real rows too, and pipelines load them.

Without relation rows the nodes stay apart. The objects exist, but nothing connects them, so widening the range turns up nowhere to go. When an expected connection is missing, the cause is usually one of three.

  • The relation rows have not been loaded yet.
  • The values passed to the relation's reference columns do not match the identifier key values of the entities.
  • The connection is several steps away and lies outside the range currently in view.

A graph is also the place to confirm whether data is genuinely connected. A correct model with a mismatched load shows that fact plainly.

Questions worth asking as a graph

Use a graph when the connection itself is the answer: what affects what, how far an object reaches, whether a path exists between two objects.

Questions that aggregate values — sums, averages, rankings — belong to tables and dashboards. A graph can count them, but using it that way is holding the wrong tool. Checking whether the answer requires following connections makes the choice easy.