01 · What we are building
System Overview
Data Services is a five-plane platform: an edge plane on the screen, an ingest plane, a processing plane, a warehouse plane, and a serving plane. Each plane is independently deployable and independently scalable, because the screen fleet grows 20x in 24 months while the analytics workload grows with client count, not screen count.
End-to-end data flow: capture → clean → map → house → analyse → publish
Every event in the platform travels this path. Nothing reaches a client surface without passing the consent filter, the quality gate and the sample threshold check.
01
Capture (edge)
Android media player on each screen; offline-first.
- Proof-of-play, screen ID, creative ID, dwell
- GPS fix at play start/end, speed, heading
- Touch events, QR token issue, survey answers
- Local SQLite buffer, store-and-forward on 4G
02
Ingest
Signed batch upload to an HTTPS collector.
- Device auth via per-device key + rotating token
- Idempotency key per event (screen+play+ts)
- Raw landing zone in object storage (immutable)
- Dead-letter queue for malformed payloads
03
Clean & validate
Schema, quality and consent enforcement.
- Schema contract validation (reject or quarantine)
- Dedup, clock-skew correction, GPS jitter smoothing
- Survey fraud checks: speeding, straight-lining, dupes
- Consent-tier filter applied before any PII persists
04
Map & enrich
Turn raw coordinates into meaning.
- Map-match GPS to road network and trip legs
- Reverse-geocode to district / subdistrict
- POI and mall proximity, daypart, traffic band
- Identity resolution: passenger key ↔ QR ↔ client feed
05
House
Warehouse with a medallion layout.
- Bronze: raw immutable events
- Silver: cleaned, conformed, enriched fact tables
- Gold: campaign, panel and benchmark marts
- Partitioned by date + region, PII in a vaulted schema
06
Analyse
Metrics, weighting and causal models.
- Semantic metric layer: one definition per metric
- Panel weighting to a Bangkok population frame
- Geo-holdout and incrementality models
- MotionReach Conversion Index computation
07
Publish
Only what passes the threshold gate.
- Client dashboard + scheduled PDF/CSV exports
- Syndicated benchmark cuts (k-anonymity enforced)
- Audience segments pushed to the ad server / SSP
- Partner API with per-tenant scoping and rate limits
02 · The edge is the hardest part
Capture Layer & Device Contract
Taxis lose connectivity constantly. The device is therefore the source of truth for a bounded window: it must record everything locally, survive power cycles mid-trip, and reconcile on reconnect without duplicating or losing plays. All device software ships behind staged OTA with automatic rollback.
On-device capture pipeline
Runs entirely offline; upload is opportunistic.
01
Playout
Scheduler plays creative from a local cache.
- Play start/end timestamps
- Creative + campaign IDs
- Screen and vehicle ID
02
Sensors
GPS sampled at 1 Hz, aggregated per play.
- Fix at start, end, midpoint
- Speed / stationary detection
- Trip leg stitching
03
Interaction
Consent, survey and QR UI.
- Consent tier state machine
- Micro-survey responses
- Per-play signed QR token
04
Buffer
Encrypted SQLite write-ahead store.
- 7-day local retention
- Backpressure and compaction
- Integrity checksum per batch
05
Sync
Batched upload with exponential backoff.
- gzip + signed payload
- Idempotent replay
- Ack-then-delete semantics
| Event | Emitted when | Key fields | Volume at 10k screens |
|---|---|---|---|
| play.completed | Creative finishes | play_id, screen_id, creative_id, dwell_ms, geo | ~24M / day |
| trip.leg | Passenger boards/alights | trip_id, origin_h3, dest_h3, duration, distance | ~600k / day |
| consent.changed | Opt-in tier set or withdrawn | passenger_key, tier, ts, proof_hash | ~120k / day |
| survey.answer | Question answered | survey_id, q_id, answer, latency_ms | ~250k / day |
| qr.issued / qr.scanned | Token rendered / scanned | token, play_id, campaign_id, ts | ~24M / ~200k |
| conversion.matched | Client feed reconciles | token or loyalty_id, value, channel | client-dependent |
03 · Where it runs
Reference Stack & Hosting
Recommendation: cloud-native, managed-first, single primary region in Singapore (ap-southeast-1) with data residency controls for Thai personal data, and a Thai-region object store for raw PII if the PDPA position tightens. Avoid self-managed Kafka/Spark clusters until volume justifies the headcount.
| Plane | Recommended | Why | Alternative |
|---|---|---|---|
| Edge runtime | Android (Kotlin) player + encrypted SQLite | Existing hardware, offline-first | React Native shell |
| Ingest | Managed HTTPS collector + Kinesis/PubSub | Elastic, no cluster ops | Kafka (MSK) at >100k eps |
| Object storage | S3 (raw, bronze) with lifecycle rules | Cheap immutable landing zone | GCS |
| Processing | dbt + Airflow (batch), Flink-lite for streaming KPIs | SQL-first, analyst-ownable | Spark for heavy geo joins |
| Warehouse | Snowflake or BigQuery | Separation of storage/compute, geo functions | ClickHouse for cost at scale |
| Geo | H3 indexing + OSM map-matching (Valhalla) | Cheap spatial joins, no vendor lock | Google Roads API |
| Serving | Postgres read models + cached API | Fast dashboards without warehouse cost per view | Cube.js semantic API |
| App layer | TanStack Start on edge, typed server functions | Same stack as this demo | n/a |
| Identity/PII | Vault schema + tokenisation service | PII never leaves the vault; analytics uses tokens | Managed KMS + HSM |
| Observability | OpenTelemetry, Grafana, data-quality tests in dbt | One trace across device → dashboard | Monte Carlo |
Hosting topology
Fleet, platform and consumers, with the PII vault isolated from the analytics plane.
01
Fleet
10,000 screens, 4G, OTA-managed.
- Device registry + key rotation
- Staged rollout rings
- Remote health beacon
02
Edge / API tier
Stateless, autoscaled, WAF-fronted.
- Collector endpoints
- Client + partner API
- Webhook receivers for client feeds
03
Data platform
Private VPC, no public egress.
- Lake + warehouse
- Orchestration and dbt
- PII vault, separate keys and IAM
04
Consumers
Everything is a tenant-scoped read.
- Brand dashboard
- Internal network ops view
- Ad server / SSP segments
- Syndicated report engine
04 · One definition per metric
Data Model & Semantic Layer
The warehouse is a star schema around a single grain: the play. Every other fact rolls up to or joins from it. Metrics are defined once in the semantic layer and consumed identically by the dashboard, the report engine and the API, a client and an internal analyst can never see two different numbers for the same thing.
Core gold-layer entities
fact_play (play_id, screen_id, trip_id, creative_id, campaign_id, ts, dwell_ms, h3_9) fact_engagement (play_id, type[touch|scan|survey], ts, latency_ms) fact_conversion (token|loyalty_id, campaign_id, ts, channel, value_thb, match_method) dim_screen (screen_id, vehicle_id, fleet_partner, install_date, status) dim_trip (trip_id, origin_h3, dest_h3, start_ts, duration_s, purpose_inferred) dim_passenger (passenger_key, consent_tier, profile_attrs..., weight) -- token only dim_creative (creative_id, campaign_id, advertiser, category, format, duration_s) vault_identity (passenger_key -> hashed contact / loyalty ref) -- isolated schema
- Grain discipline: no metric is defined on a table whose grain it cannot be summed over.
- Slowly-changing dimensions (type 2) for screens and creatives so historic campaigns re-run identically.
- Every gold table carries lineage columns: source batch, pipeline version, quality score.
- Reproducibility: any published figure can be regenerated from bronze with a pinned pipeline version.
05 · Nothing ships unqualified
Quality, Consent & Publication Gates
Gate sequence before any number is published
A cut that fails any gate is suppressed with an explicit reason code, never silently rounded.
01
Consent gate
Tier governs field availability.
- Withdrawal propagates in <24h
- Consent proof stored per record
- Tier downgrade re-filters marts
02
Quality gate
dbt tests on every run.
- Freshness, uniqueness, referential integrity
- GPS plausibility, device clock drift
- Survey fraud scoring
03
Sample gate
Publishability thresholds.
- Min profiled rides per cut
- k-anonymity ≥ 30 for syndicated cells
- Confidence interval attached to every figure
04
Release
Versioned, auditable outputs.
- Snapshot ID on every export
- Restatement policy and changelog
- Tenant-scoped access log
06 · What it takes to build
Delivery Plan, Team & Cost
| Phase | Engineering scope | Team | Exit criteria |
|---|---|---|---|
| P0 · 0–3m | Event schema, collector, bronze lake, device SDK v1, consent UI | 1 platform eng, 1 Android eng, 1 analytics eng | 500 screens emitting validated events at >98% delivery |
| P1 · 3–6m | Warehouse + dbt silver/gold, geo enrichment, QR service, dashboard v1 | +1 data eng, +1 full-stack | First paid attribution pilot reported end-to-end |
| P2 · 6–12m | Semantic layer, weighting, benchmark engine, partner API, SSP segment push | +1 data scientist, +1 backend | Syndicated report generated automatically each month |
| P3 · 12–24m | Streaming KPIs, multi-tenant self-serve, Index validation, SEA multi-region | +1 SRE, +1 analytics eng | 10k screens supported with flat cost per screen |
- Run rate at 10k screens: expect warehouse + storage to dominate; budget cost-per-1k-plays as a tracked SLO.
- Buy over build for map-matching, geocoding and identity tokenisation; build the semantic layer and the Index in-house. That is the IP.
- Key risks: device connectivity variance, GPS accuracy in dense Bangkok, client conversion-feed latency, and warehouse cost drift.
