본문으로 건너뛰기
← Data Engineer Path

Load an API response into a dataset with a code node

12 min

What you will learn

Build the first pipeline and bind the REST connector to a code node to ingest 72 hourly readings into a dataset.

A pipeline is a workflow you assemble on a canvas from a collection's datasets, code assets, and transform nodes.

This lesson builds the first one. It has no input dataset: the thing producing data is an external API, so a code node with a connector bound to it is the pipeline's starting point.

Why this needs code

The response you inspected in lesson 02 was a nested object of column-wise arrays.

hourly.time[]                  ["2026-08-03T00:00", "2026-08-03T01:00", ...]
hourly.temperature_2m[]        [27.6, 27.1, ...]
hourly.relative_humidity_2m[]  [87, 89, ...]

A dataset is a set of rows, so those three arrays have to be zipped by index into one row each. No standard transform — rename, cast, aggregate — expresses that. Hence a code node.

Create the result dataset first

Before opening the pipeline editor, create the src_weather_hourly table in your collection. Output tables created through Quick Add can fail to resolve when the pipeline is saved.

  1. Select Collections in the left sidebar and open the training collection.
  2. Select Add item, hover over Dataset, and choose Table.
  3. Under Basic information, name it src_weather_hourly. Leave Alias blank or enter src_weather_hourly so the collection list shows the same identifier, then select Next.
Korean dataset creation form with src_weather_hourly entered as the system name
Set the system identifier that the pipeline code will reference
  1. Define four columns in Schema:
ColumnData typeMeaning
observed_atTextObservation timestamp, e.g. 2026-08-03T00:00
observed_dateTextThe date portion only. Lesson 04's grouping key
temperature_2mDoubleTemperature in °C
relative_humidity_2mBigintRelative humidity in %
  1. Replace the default field1 and add the other three columns. Once the fields are real, the bottom action changes to Create; select it.
Korean schema form with observed_at, observed_date, temperature_2m, and relative_humidity_2m
Define two Text columns, one Double column, and one Bigint column

Leaving two column names exactly as the API returned them is deliberate. The raw layer preserves the API's vocabulary; readable names are the next lesson's job.

Korean engineer training collection showing the newly created src_weather_hourly dataset
Confirm src_weather_hourly appears in the collection after creation

Create the Python code asset

Create the code as a collection asset before opening the pipeline. A temporary code node made through Quick Add may not resolve to a valid code ID when the pipeline is saved, so this course does not use that route.

  1. In the training collection, select Add item → Code → Python.
  2. Set both name and alias to collect_weather_hourly.
  3. Enter Collect hourly weather from Open-Meteo as the description.
  4. Select Next. The code below uses the runtime's built-in polars, so no additional package declaration is needed.

Bind the connector

Using a connector from code requires a @use_connector declaration. Let the editor insert it rather than typing it.

  1. Type @ on the first line of the code editor.
  2. Choose connector from the menu that appears.
  3. Choose open_meteo from the connector list.

This line is inserted:

@use_connector("open_meteo", ref="open_meteo")

The first argument is the variable name your code uses; ref is the connector name to look up. At run time the runtime resolves that name, builds the connection object, and injects it into the variable. The address and credentials stored on the connector never appear in your code.

Write the code

Below the decorator, write the following code and select Create:

@use_connector("open_meteo", ref="open_meteo")
def run(options=None, contexts=None):
    import polars as pl

    body = open_meteo.get(  # noqa: F821 — injected by @use_connector
        "/v1/forecast",
        {
            "latitude": 37.5665,
            "longitude": 126.9780,
            "hourly": "temperature_2m,relative_humidity_2m",
            "past_days": 3,
            "forecast_days": 0,
            "timezone": "Asia/Seoul",
        },
    )

    hourly = body["hourly"]
    frame = pl.DataFrame(
        {
            "observed_at": hourly["time"],
            "observed_date": [value[:10] for value in hourly["time"]],
            "temperature_2m": hourly["temperature_2m"],
            "relative_humidity_2m": hourly["relative_humidity_2m"],
        },
        schema_overrides={
            "observed_at": pl.String,
            "observed_date": pl.String,
            "temperature_2m": pl.Float64,
            "relative_humidity_2m": pl.Int64,
        },
    )
    return {"src_weather_hourly": frame}
Korean code editor showing the Open-Meteo request and Polars transformation
Paste the Python code without changing its indentation

After creation, open collect_weather_hourly from the collection and confirm that the source appears in the Code tab. If only the asset name exists and the editor is empty, choose Edit → Code, paste it again, and select Save changes. A code asset without a source file fails during pipeline save or startup.

Three things to notice:

  • run takes no input argument. A code node with no input dataset connected is called without one. The connector is the only source of data.
  • get returns parsed JSON. You get a dictionary, not a response object, so body["hourly"] works directly. The path is appended to the connector's Base URL.
  • The returned dictionary's key is the output port name. src_weather_hourly must match the output dataset you connect in a moment.

observed_date comes from slicing the first ten characters of the timestamp. Lesson 04's aggregation groups on that column.

The returned frame is a Polars DataFrame. schema_overrides sets the four dataset column types, keeping them stable for an empty response or null readings. The next lesson's code nodes also receive table inputs as Polars DataFrames.

Open the editor and pick a collection

  1. Select Data → Pipelines in the left sidebar.
  2. Select Create in the upper-right.
  3. Choose the training collection on Select pipeline collection. The editor opens as soon as you choose it; the current UI does not require a separate Continue click.
Korean screen for choosing the collection that owns a new pipeline
Choose the collection containing the dataset and Python code

You cannot add items or run anything before choosing a collection, and a pipeline can only use assets from the collection you chose.

Connect the output dataset

  1. Search for collect_weather_hourly in the Component Library. Drag the vertical-dot handle on the item's right edge onto an empty part of the canvas.
  2. Search for src_weather_hourly and drag it with the same right-edge handle. Dragging the name itself may only select the item.
  3. Connect the code node's right handle to the dataset's left handle.
Korean pipeline canvas with collect_weather_hourly connected to src_weather_hourly
Connect the code output handle to the dataset input handle
  1. Select the code node and open the inspector's Options tab.
  2. Expand the src_weather_hourly row under Output (1).
  3. Confirm the return key and dataset, change Write mode to Overwrite, and select Save at the bottom of the inspector.
Korean code node options showing src_weather_hourly and Overwrite write mode
Expand the output row and change its write mode to Overwrite
collect_weather_hourly → src_weather_hourly

If the output port name differs from the returned dictionary's key, the run fails. Compare the name in Output against the return {"...": frame} key side by side.

Save and run once

  1. Select Save in the upper-right.
  2. Name it weather_daily_pipeline and leave the type as Batch.
  3. Once saved, select Run now.
Korean pipeline editor showing a successful weather pipeline run
Confirm the run status changes to Success and the completion notice appears

You grow this same pipeline through the remaining lessons rather than creating a new one each time.

After the run, open src_weather_hourly in the collection and check the Data tab:

  • Are there 72 rows? That's 24 hours × 3 days.
  • Does observed_date hold three distinct dates?
  • Does temperature_2m carry decimals and relative_humidity_2m whole numbers?
Korean src_weather_hourly Data tab showing 72 rows and four weather columns
Verify 72 rows and the values in all four weather columns

Check yourself

  • Did you create src_weather_hourly with its four columns before building the pipeline?
  • Does the code node declare @use_connector("open_meteo", ref="open_meteo")?
  • Are the returned dictionary key and the output port name both src_weather_hourly?
  • Is the write mode Overwrite?
  • Did the run load 72 rows?

Next lesson

You attach a preparation code node to the same pipeline: rename the API's columns into readable ones, then aggregate 72 hourly rows into three daily rows for an analysis-ready mart.

Before you finish

Use these questions to check whether you achieved this lesson's goal.

  • Can you repeat ‘Load an API response into a dataset with a code node’ without following the instructions?
  • Can you name at least one place to check when the result differs from what you expected?