본문으로 건너뛰기
Data Engineer
55 min

Data Engineer Path

A 55-minute path for data engineers. Connect a public REST API, ingest from it, transform and aggregate the result, put it on a schedule, and hand it to an analyst as a dashboard — all in one pipeline.

0/6 complete

About this Path

A path for data engineers who are new to D.Hub.

You register a public weather API as a data connection, ingest its hourly readings into a dataset, aggregate them into a daily summary, put the whole thing on a daily schedule, and finish with the dashboard chart an analyst sees. After six lessons, your training collection holds one working pipeline and one dashboard built on its output.

The point of this path is that everything from attaching an external system to handing the result to someone else forms a single line.

Loading the diagram. Mermaid source:

flowchart LR
    accTitle: Data Engineer path end-to-end flow
    accDescr: A public REST API is registered as a data connection, ingested by a code node into an hourly dataset, cleaned and aggregated with Polars into a daily mart, put on a schedule with run monitoring, and handed off as a dashboard chart.
    api[(Public weather API)] --> connector[REST data connection]
    connector --> code[Code node ingest]
    code --> src[(Hourly readings)]
    src --> prepare[Preparation code]
    prepare --> stg[(Cleaned hourly)]
    prepare --> mart[(Daily summary mart)]
    schedule[Daily schedule] --> code
    mart --> dashboard[Dashboard chart]

Terms to know before you start

  • Data connection (connector): A reusable bundle of address and authentication settings for an external system. This path registers a public weather API here.
  • Pipeline: An execution flow that reads data, transforms it, and writes it somewhere. This path grows one pipeline lesson by lesson.
  • Code node: A pipeline step written in Python for work no standard transform can express. Calling an external API is the classic case.
  • Transform node: A no-code step for renaming, casting, filtering, joining, and aggregating.
  • Mart: The result dataset an analyst can use directly, once aggregation and naming are settled.

Prerequisites

  • Access to the D.Hub portal. Registering a data connection requires the admin or manager role.
  • One collection for the exercise. Every connector, dataset, code asset, pipeline, and dashboard in this path lands there.
  • Either the Analyst Path or a working understanding of collections and datasets.
  • A D.Hub environment with outbound internet access. The connection test runs from the manager process, not your browser, so a blocked outbound path stops you in lesson 02.

Nothing here requires a signup or an API key. The API this path uses is a public service you can call without authentication.

Prior experience with dbt, Airflow, or Snowflake will help you map D.Hub's concepts faster.

What you'll be able to do

  • Register a managed REST data connection and inspect its response with the connection test and Query Console.
  • Bind a data connection to a code node and load an external API's response into a dataset.
  • Clean a raw dataset and build an analysis-ready daily mart with a Polars code node.
  • Choose write and read modes so repeated runs do not corrupt the result.
  • Put a batch pipeline on a recurring schedule and read its run history.
  • Trace a failed run to its cause in history, fix it, and rerun the pipeline.
  • Document a result dataset, set permissions, and hand it to an analyst as a dashboard chart.

What comes after this Path

Selecting the checkbox next to each lesson saves your progress automatically.

Lessons

  1. 01Data engineer's main work areasThe four surfaces engineers use in portal — Connectors, Pipelines, Codes, Datasets — mapped against dbt and Airflow, plus what this path builds.
    7 min
  2. 02Register a public REST API as a data connectionRegister a public weather API as a managed REST connector, then confirm what it actually returns with the connection test and the Query Console.
    10 min
  3. 03Load an API response into a dataset with a code nodeBuild the first pipeline and bind the REST connector to a code node to ingest 72 hourly readings into a dataset.
    12 min
  4. 04Clean the raw data and build a daily mart with codeUse Polars to rename the API columns and summarize 72 hourly rows into a three-row daily mart.
    12 min
  5. 05Run a batch pipeline on a scheduleAdd a recurring schedule to the ingest-to-mart pipeline and see why running it every day does not corrupt the result.
    8 min
  6. 06Diagnose a failure and hand off through a dashboardTrace a code error in run history and fix it, then put the daily mart on a dashboard chart and settle its description and permissions for an analyst.
    12 min