v1.0 · August 2026

Data Services Technical Specification

Reference architecture, data pipeline and platform plan, for CTO review

Prepared for the MotionReach CTO and engineering leadership

01 · What we are building

System Overview

Data Services is a five-plane platform: an edge plane on the screen, an ingest plane, a processing plane, a warehouse plane, and a serving plane. Each plane is independently deployable and independently scalable, because the screen fleet grows 20x in 24 months while the analytics workload grows with client count, not screen count.

End-to-end data flow: capture → clean → map → house → analyse → publish

Every event in the platform travels this path. Nothing reaches a client surface without passing the consent filter, the quality gate and the sample threshold check.

01

Capture (edge)

Android media player on each screen; offline-first.

  • Proof-of-play, screen ID, creative ID, dwell
  • GPS fix at play start/end, speed, heading
  • Touch events, QR token issue, survey answers
  • Local SQLite buffer, store-and-forward on 4G

02

Ingest

Signed batch upload to an HTTPS collector.

  • Device auth via per-device key + rotating token
  • Idempotency key per event (screen+play+ts)
  • Raw landing zone in object storage (immutable)
  • Dead-letter queue for malformed payloads

03

Clean & validate

Schema, quality and consent enforcement.

  • Schema contract validation (reject or quarantine)
  • Dedup, clock-skew correction, GPS jitter smoothing
  • Survey fraud checks: speeding, straight-lining, dupes
  • Consent-tier filter applied before any PII persists

04

Map & enrich

Turn raw coordinates into meaning.

  • Map-match GPS to road network and trip legs
  • Reverse-geocode to district / subdistrict
  • POI and mall proximity, daypart, traffic band
  • Identity resolution: passenger key ↔ QR ↔ client feed

05

House

Warehouse with a medallion layout.

  • Bronze: raw immutable events
  • Silver: cleaned, conformed, enriched fact tables
  • Gold: campaign, panel and benchmark marts
  • Partitioned by date + region, PII in a vaulted schema

06

Analyse

Metrics, weighting and causal models.

  • Semantic metric layer: one definition per metric
  • Panel weighting to a Bangkok population frame
  • Geo-holdout and incrementality models
  • MotionReach Conversion Index computation

07

Publish

Only what passes the threshold gate.

  • Client dashboard + scheduled PDF/CSV exports
  • Syndicated benchmark cuts (k-anonymity enforced)
  • Audience segments pushed to the ad server / SSP
  • Partner API with per-tenant scoping and rate limits

02 · The edge is the hardest part

Capture Layer & Device Contract

Taxis lose connectivity constantly. The device is therefore the source of truth for a bounded window: it must record everything locally, survive power cycles mid-trip, and reconcile on reconnect without duplicating or losing plays. All device software ships behind staged OTA with automatic rollback.

On-device capture pipeline

Runs entirely offline; upload is opportunistic.

01

Playout

Scheduler plays creative from a local cache.

  • Play start/end timestamps
  • Creative + campaign IDs
  • Screen and vehicle ID

02

Sensors

GPS sampled at 1 Hz, aggregated per play.

  • Fix at start, end, midpoint
  • Speed / stationary detection
  • Trip leg stitching

03

Interaction

Consent, survey and QR UI.

  • Consent tier state machine
  • Micro-survey responses
  • Per-play signed QR token

04

Buffer

Encrypted SQLite write-ahead store.

  • 7-day local retention
  • Backpressure and compaction
  • Integrity checksum per batch

05

Sync

Batched upload with exponential backoff.

  • gzip + signed payload
  • Idempotent replay
  • Ack-then-delete semantics
EventEmitted whenKey fieldsVolume at 10k screens
play.completedCreative finishesplay_id, screen_id, creative_id, dwell_ms, geo~24M / day
trip.legPassenger boards/alightstrip_id, origin_h3, dest_h3, duration, distance~600k / day
consent.changedOpt-in tier set or withdrawnpassenger_key, tier, ts, proof_hash~120k / day
survey.answerQuestion answeredsurvey_id, q_id, answer, latency_ms~250k / day
qr.issued / qr.scannedToken rendered / scannedtoken, play_id, campaign_id, ts~24M / ~200k
conversion.matchedClient feed reconcilestoken or loyalty_id, value, channelclient-dependent

03 · Where it runs

Reference Stack & Hosting

Recommendation: cloud-native, managed-first, single primary region in Singapore (ap-southeast-1) with data residency controls for Thai personal data, and a Thai-region object store for raw PII if the PDPA position tightens. Avoid self-managed Kafka/Spark clusters until volume justifies the headcount.

PlaneRecommendedWhyAlternative
Edge runtimeAndroid (Kotlin) player + encrypted SQLiteExisting hardware, offline-firstReact Native shell
IngestManaged HTTPS collector + Kinesis/PubSubElastic, no cluster opsKafka (MSK) at >100k eps
Object storageS3 (raw, bronze) with lifecycle rulesCheap immutable landing zoneGCS
Processingdbt + Airflow (batch), Flink-lite for streaming KPIsSQL-first, analyst-ownableSpark for heavy geo joins
WarehouseSnowflake or BigQuerySeparation of storage/compute, geo functionsClickHouse for cost at scale
GeoH3 indexing + OSM map-matching (Valhalla)Cheap spatial joins, no vendor lockGoogle Roads API
ServingPostgres read models + cached APIFast dashboards without warehouse cost per viewCube.js semantic API
App layerTanStack Start on edge, typed server functionsSame stack as this demon/a
Identity/PIIVault schema + tokenisation servicePII never leaves the vault; analytics uses tokensManaged KMS + HSM
ObservabilityOpenTelemetry, Grafana, data-quality tests in dbtOne trace across device → dashboardMonte Carlo

Hosting topology

Fleet, platform and consumers, with the PII vault isolated from the analytics plane.

01

Fleet

10,000 screens, 4G, OTA-managed.

  • Device registry + key rotation
  • Staged rollout rings
  • Remote health beacon

02

Edge / API tier

Stateless, autoscaled, WAF-fronted.

  • Collector endpoints
  • Client + partner API
  • Webhook receivers for client feeds

03

Data platform

Private VPC, no public egress.

  • Lake + warehouse
  • Orchestration and dbt
  • PII vault, separate keys and IAM

04

Consumers

Everything is a tenant-scoped read.

  • Brand dashboard
  • Internal network ops view
  • Ad server / SSP segments
  • Syndicated report engine

04 · One definition per metric

Data Model & Semantic Layer

The warehouse is a star schema around a single grain: the play. Every other fact rolls up to or joins from it. Metrics are defined once in the semantic layer and consumed identically by the dashboard, the report engine and the API, a client and an internal analyst can never see two different numbers for the same thing.

Core gold-layer entities

fact_play           (play_id, screen_id, trip_id, creative_id, campaign_id, ts, dwell_ms, h3_9)
fact_engagement     (play_id, type[touch|scan|survey], ts, latency_ms)
fact_conversion     (token|loyalty_id, campaign_id, ts, channel, value_thb, match_method)
dim_screen          (screen_id, vehicle_id, fleet_partner, install_date, status)
dim_trip            (trip_id, origin_h3, dest_h3, start_ts, duration_s, purpose_inferred)
dim_passenger       (passenger_key, consent_tier, profile_attrs..., weight)   -- token only
dim_creative        (creative_id, campaign_id, advertiser, category, format, duration_s)
vault_identity      (passenger_key -> hashed contact / loyalty ref)          -- isolated schema
  • Grain discipline: no metric is defined on a table whose grain it cannot be summed over.
  • Slowly-changing dimensions (type 2) for screens and creatives so historic campaigns re-run identically.
  • Every gold table carries lineage columns: source batch, pipeline version, quality score.
  • Reproducibility: any published figure can be regenerated from bronze with a pinned pipeline version.

05 · Nothing ships unqualified

Quality, Consent & Publication Gates

Gate sequence before any number is published

A cut that fails any gate is suppressed with an explicit reason code, never silently rounded.

01

Consent gate

Tier governs field availability.

  • Withdrawal propagates in <24h
  • Consent proof stored per record
  • Tier downgrade re-filters marts

02

Quality gate

dbt tests on every run.

  • Freshness, uniqueness, referential integrity
  • GPS plausibility, device clock drift
  • Survey fraud scoring

03

Sample gate

Publishability thresholds.

  • Min profiled rides per cut
  • k-anonymity ≥ 30 for syndicated cells
  • Confidence interval attached to every figure

04

Release

Versioned, auditable outputs.

  • Snapshot ID on every export
  • Restatement policy and changelog
  • Tenant-scoped access log

06 · What it takes to build

Delivery Plan, Team & Cost

PhaseEngineering scopeTeamExit criteria
P0 · 0–3mEvent schema, collector, bronze lake, device SDK v1, consent UI1 platform eng, 1 Android eng, 1 analytics eng500 screens emitting validated events at >98% delivery
P1 · 3–6mWarehouse + dbt silver/gold, geo enrichment, QR service, dashboard v1+1 data eng, +1 full-stackFirst paid attribution pilot reported end-to-end
P2 · 6–12mSemantic layer, weighting, benchmark engine, partner API, SSP segment push+1 data scientist, +1 backendSyndicated report generated automatically each month
P3 · 12–24mStreaming KPIs, multi-tenant self-serve, Index validation, SEA multi-region+1 SRE, +1 analytics eng10k screens supported with flat cost per screen
  • Run rate at 10k screens: expect warehouse + storage to dominate; budget cost-per-1k-plays as a tracked SLO.
  • Buy over build for map-matching, geocoding and identity tokenisation; build the semantic layer and the Index in-house. That is the IP.
  • Key risks: device connectivity variance, GPS accuracy in dense Bangkok, client conversion-feed latency, and warehouse cost drift.