urbanpulse
Enter the cityEnter
Explore UrbanPulse

A different way to read the city.

Enter the city Public signals. A wider perspective.
URBANPULSE / ENGINEERING CASE STUDY

The last kilometre
of data engineering.

Reliable pipelines are only the beginning. The real work is turning fragmented data into something people can understand and use.

THE PROBLEM

A city’s weather, air, and mobility data live in different systems. Each comes with its own schema, clock, geography, and definition of “current”.

THE APPROACH

Retain the originals. Normalize at the boundary. Model useful questions. Carry provenance and quality all the way to the interface.

One system. Four deliberate boundaries.

Select a stage to understand what it does and why it exists.

01 / OPEN-METEO / CAMS / GBFS / OPENAQ

Different schemas, timestamps, units and refresh rates. Sources are fetched by server-side adapters; OpenAQ is an optional keyed provider. Casablanca intentionally exposes partial coverage.

Model the questions, not the providers.

ANALYTICAL MODEL

fct_weather_hourly

City, time, temperature, humidity, precipitation and wind. Weather-model observations retain their valid time.

ANALYTICAL MODEL

fct_air_quality_measurements

Location, pollutant, source unit, valid time and modeled/measured context. No automatic unit mixing or health labeling.

ANALYTICAL MODEL

fct_mobility_station_snapshots

Station, collection time, capacity, available bikes and docks. An inventory snapshot is not a trip counter.

ANALYTICAL MODEL

mart_city_hourly

Hourly city signals joined at matching UTC boundaries. The API receives bounded, useful aggregates.

Reliability should be visible.

Keep the evidence. Original payloads are retained before parsing, so a malformed source response remains available for investigation. A failed live provider stays failed; it never turns into a fixture.

Expect duplicates. The event boundary uses at-least-once delivery. Content IDs and unique observation keys make processing repeatable. Consumer offsets advance after the database commit.

Preserve uncertainty. Missing readings remain null. Freshness thresholds belong to each source. The UI distinguishes unavailable, stale and current signals.

Make time explicit. All observations are stored in UTC. City-local time is presentation only. Replay applies an “at or before” cutoff independently to each signal.

Trade-offs, made explicit.

Why Kafka and batch coexist
Mobility snapshots can arrive independently of downstream processing. Kafka provides a durable replay boundary. Airflow schedules slower environmental collection and dbt builds.
Why PostgreSQL / PostGIS
One store handles integrity, time windows and spatial queries. The workload does not justify an additional analytical database.
Why retain raw JSON
Source documents are small and nested. JSON preserves the provider response faithfully. Columnar history can be introduced when a measured workload warrants it.
Why dbt
Transformations, tests and descriptions are version controlled. The manifest gives the product real lineage instead of a manually maintained diagram.
Why scheduled collection
Weather and air are collected every 30 minutes, and mobility every 5 minutes. Replay grows from retained observations. Source failures remain visible while the next collection retries independently.
BUILT BY YASSINE ERRADOUANI

Data engineering, all the way to the product.

Python · Kafka · Airflow · dbt · PostgreSQL / PostGIS · FastAPI · Next.js · TypeScript