urbanpulse
Enter the cityEnter
Explore UrbanPulse

A different way to read the city.

Enter the city Public signals. A wider perspective.
Project journal01 / UrbanPulse

Following the signals: building UrbanPulse

My projects tend to take me into different fields. With UrbanPulse, I wanted to build something distinct and put my own touch on how it looked and worked. The idea became a city view that brings weather, air quality and bike availability together.

I invited Wissal Charkaoui to help with the engineering side. She was interested in data engineering and seemed like a good partner. Having a helping hand would give me more room for the heavier work and for polishing the application as a whole.

The application covers Paris and Casablanca, but their coverage isn't identical. Building the city view also meant keeping track of where each number came from and when it was actually reported. Those details became part of the product.

Starting with the feeds

The weather adapter uses Open-Meteo's current and hourly data. Air quality comes from CAMS through Open-Meteo's air-quality API. Paris mobility comes from the Vélib' GBFS feed, which separates station information from station status.

These sources describe different things. CAMS provides modeled air conditions at a grid location; it isn't a street-level sensor reading. A bike count belongs to a station snapshot. Putting the values beside each other doesn't remove those differences, so the interface keeps their units, source names and timestamps available.

The GBFS adapter starts from the discovery endpoint and finds the station feeds there. That keeps provider URLs out of the rest of the application. The adapters normalize their payloads into weather, air and mobility records before the API reads them.

Keeping the originals

The ingestion path saves raw JSON before writing normalized observations. Source, city, mode and payload go into a SHA-256 content ID. If the same event arrives again, that ID and the observation keys keep repeated delivery from creating duplicate records.

The raw payload is useful when a normalized value needs an explanation. It leaves something concrete to inspect alongside the pipeline run and quality results. Retries are bounded, and unrecoverable collection failures are recorded instead of being replaced with a successful-looking response.

I kept PostgreSQL with PostGIS as the production store. The application needs time-window queries, station positions and hourly aggregates. The storage decision records a separate analytical database or geospatial service as extra deployment and synchronization work without a measured need for it yet.

SQLite remains available for local development and tests. It's convenient when Docker isn't running, but it doesn't provide the PostGIS part of the platform. The production settings require PostgreSQL and live mode.

Replay follows the same rule about keeping the evidence. It reads stored observations with a time cutoff. Its history grows as the collector saves station snapshots; it doesn't invent frames to fill gaps.

Keeping the first deployment small

The repository includes Kafka, Airflow and dbt, but they aren't all required by the production Compose file. The optional platform profile gives frequent mobility collection a Kafka boundary, lets Airflow schedule the work, and uses dbt for versioned analytical models.

That boundary has a cost. At-least-once delivery means the consumer can see an event more than once, so persistence has to remain idempotent. The consumer commits its offset after normalization and the quality pass have committed. Streaming the slower weather and air feeds would add machinery without the same benefit.

For the first Oracle deployment, the smaller stack uses a direct Python collector, FastAPI, PostgreSQL/PostGIS, Next.js and Caddy. Weather and air collection run every thirty minutes; Paris mobility runs every five. The optional platform remains available without becoming a requirement for serving the site.

The dbt switch is explicit too. The API only reads the hourly mart when DBT_ENABLED is enabled; the small production stack leaves it disabled and reads normalized data. That keeps a partially configured database from pretending a model has been built.

I'm preparing this deployment around an Oracle Always Free VM. The Docker setup and deployment guide are ready, but preparing a deployment isn't the same as having it running on Oracle. That's still the next step.

Not quite that simple

One actual feed mismatch was the timestamp in Vélib's status document. It uses lastUpdatedOther, while the adapter also needs to handle the usual GBFS last_updated field. The adapter accepts either and falls back to the newest station last_reported value if neither is present.

That matters because collection time and observation time aren't interchangeable. A successful request can retrieve an old snapshot. Source freshness uses a source-reported time, and the retained payload keeps the individual station report times available for investigation.

Casablanca has a different limitation. Weather and modeled air are configured, but no public GBFS feed has been verified for it. Its mobility metric, station layer and replay therefore stay unavailable. The application doesn't borrow Paris data or fill the missing domain with demo values.

Demo fixtures still have a place in development and tests. They run through the normalization path, are labeled in the interface, and stay separate from live rows by mode. Production rejects demo mode. A stale or missing live source remains a stale or missing live source.

Making the interface feel like a place

The frontend uses Next.js, React and TypeScript, with MapLibre for the city maps. Getting the data onto a screen wasn't the whole design problem. The mobile version felt too much like elements stacked on top of each other, and I wanted the background to have more room.

The hero copy now sits toward the bottom of the phone screen. Weather, environment and mobility use a horizontal signal strip on mobile, while the larger sections reveal in sequence as you move down the page. GSAP and ScrollTrigger handle those section transitions.

The smaller interactions need their own rules. The cards beneath the surface reveal their explanations on hover on desktop. Touch screens keep a tap-to-expand version, and keyboard users can reach the explanations too. The behavior follows the way the device is used instead of relying on a hover state everywhere.

Source context is still part of that interface. The quality and pipeline screens expose missing fields, freshness and collection runs. A map or animation can make the site easier to read, but it doesn't make an observation more current.

Where it goes from here

There are still clear limits to this version. Casablanca's mobility coverage is incomplete, modeled air isn't a local sensor measurement, and the default seven-day retention window bounds the available replay history. Those are properties of the data and the current setup, not things the interface can fix on its own.

A longer replay window would mean revisiting retention and storage. Adding another city would start with verifying its providers. Neither change needs to be hidden behind a bigger architecture before there's a concrete reason for it.

UrbanPulse now has the city view and the collection, storage and quality paths that support it. The next step is getting that live stack onto the Oracle VM.