Skip to content

Backfilling and Forwarding Data to the Collector

Overview

Audience: Customer engineers and integration partners feeding the Collector — seeding the content catalog, wiring a customer data platform (CDP), analytics stack, or homegrown event source, or both.

Compass learns from user–item interactions: it needs to know what content exists (the catalog) and what your audience does with it (the events). Producing both is a job on your side. This guide covers the two producer-side jobs:

  • Seed and keep the content catalog current — bulk-load your existing back catalog so the model knows what items exist, then wire an ongoing live sync. See Bulk-loading the content catalog.
  • Forward behavioral events — translate your source’s events (page views, clicks, scroll depth) into the canonical event envelope and POST them to the Collector. See Forwarding events from a CDP or server-side source.

The two are complementary. A typical onboarding is: bulk-load the back catalog once, wire ongoing content sync, and forward events from your CDP or analytics source in parallel. The Compass Onboarding Checklist gives the full sequence.

Before you start: auth, endpoints, and retries

Both jobs below hit the same Collector, authenticate the same way, and share the same error-handling contract. This section states those shared rules once; each half references it rather than repeating it.

Endpoints

EndpointUse
POST /collector/v1/contentA single UpsertContent object — one catalog item per request.
POST /collector/v1/eventsA single event-envelope object.
POST /collector/v1/events/batchA JSON array of up to 100 event envelopes.

The content endpoint has no batch variant — “bulk load” means streaming items through the single-item endpoint (see Chunking and submitting a backlog). Only the events endpoint batches, at up to 100 per request.

Your tenant is resolved automatically from your unique base URL (https://{org}-config-prod.api.arc-cdn.net) — your Arc contact provides the exact host during onboarding.

Authentication

Authenticate every request with your Collector (Collector API Headless) API token in the X-API-Key header — not the Recommendations API token. It is a server-side secret: load it from a secret store, never hard-code it. The token must be provisioned against the Collector API key collection; a token from the wrong key collection fails auth. See Compass Authentication and Tokens for provisioning, key-collection separation, and handling rules.

Error and retry behavior

Every Collector POST — content or event, single or batch — returns the same status contract:

StatusMeaningCaller action
202 (or 2xx)The item/event was durably accepted for processing.None. A 202 confirms the payload was queued, not that it is fully visible to the model yet.
401 / 403X-API-Key is missing, unknown, or the wrong host.Stop. Verify the token, that it was provisioned against the Collector API key collection, and the base URL.
422The payload failed validation (bad enum, naive datetime, missing required field, unknown event_type, oversize field).Do not retry as-is — the response body identifies the offending field via detail[*].loc. Fix the row/mapping and resubmit.
5xx / 429Transient server-side or rate error.Retry the same request with exponential backoff. Cap attempts so a stuck job doesn’t stall the rest of the load.

A 422 on one item does not stop the rest — each request is independent. Surface the failed row to your operator with the response body.

A stable item_id underpins everything

Both jobs key on item_id, and it must be stable: the same content item must always produce the same item_id, and the value the bulk load sends must match what the live content sync and every forwarded event will send for that same item. There is no retroactive rebind. Settle your scheme — stability, format, length, no-PII — against the Item ID Guidance before you send the first request. The same discipline applies to user_id on events; see the User ID Guidance.

Bulk-loading the content catalog

This half explains how to seed the Compass catalog from an existing CMS export — without waiting for a live CMS webhook or IFX integration to backfill it organically.

The Collector API does not expose a dedicated batch endpoint for content. “Bulk load” means streaming each item from your export through the single-item content endpoint, POST /collector/v1/content, with appropriate pacing, idempotency, and verification. Auth, the endpoint’s status codes, and retry behavior are the shared rules in Before you start.

When to use bulk load vs IFX streaming

Use bulk load when…Use IFX streaming when…
You are onboarding to Compass and need to seed your existing back catalog.You are wiring the ongoing, live sync from your CMS to Compass.
You are re-importing after a CMS migration, a content model change, or a catalog cleanup.You want every publish, update, and unpublish to reach Compass within seconds, automatically.
You have a one-shot export (CSV / JSON / a database dump) and need to push it in.You are an Arc XP CMS customer and the IFX Recipe is already an option.

The two paths are complementary, not exclusive. The typical onboarding flow is: bulk-load the back catalog once, then wire IFX (or your own webhook) for ongoing changes. See the Compass Onboarding Checklist for the full sequence.

Prerequisites

  • Compass service provisioned.
  • A Collector API Headless API token (see Authentication).
  • An export of your published catalog. Deleted, unpublished, and draft items must not be sent.

Endpoint summary

FieldValue
Method + pathPOST /collector/v1/content
Request bodyA single JSON UpsertContent payload — one item per request.
Semanticsaction: "publish" is an upsert on (site_id, item_id). Reposting the same payload is safe.

Status codes and retry behavior are the shared contract in Before you start.

See the POST /collector/v1/content reference for the full schema, including all optional fields.

Input format and field mapping

Every row in your export maps to one `UpsertContent` payload (action: "publish"). At minimum, each payload needs:

Payload fieldWhat it isWhere it comes from in a typical CMS export
actionAlways "publish" for bulk load.Constant.
item_idStable identifier for the content item.The CMS’s primary identifier for the item (e.g. an ArcID).
site_idThe website / property the item belongs to.The site or brand the item is published under.
typeOne of article, video, podcast.The item’s content type.
timestampPublication date, ISO 8601 with explicit offset.The item’s published-at timestamp, not the export time.
titleDisplay title.The item’s headline / title field.
categories, tags, author, is_premium, metadataOptional but recommended.Whatever taxonomy and attribution fields your export carries.

Settle your item_id scheme before you start the load

A bulk load is the moment you commit to your item_id scheme — see A stable item_id underpins everything for the general rule. The one point specific to bulk loading:

  • Match what the live path will use. Whichever item_id your live CMS webhook (or IFX integration) will send for a given item after the bulk load is exactly what the bulk load must send for the same item now. A mismatch orphans the bulk-loaded record the first time that item updates through the live path.

For the complete list of content-payload fields (required and optional), see the Content endpoint section of the Compass API Developer Guide.

Chunking and submitting a backlog

The content endpoint takes one item per request, so “bulk” is a client-side concern: how many items per second you push, and how many you push in parallel.

  • Pace on aggregate throughput, not per-request. Whether you run the load as one process or many parallel workers, Compass sees the sum of your request rate, and that aggregate is the only number that matters. Per-request 202s confirm the payload was queued — they do not confirm the downstream pipeline is keeping up. Start conservative (a few tens of requests per second across all workers is a reasonable opener for an initial backfill), watch for 5xx responses and growing latency, and ramp only when you have evidence of headroom.
  • Use modest concurrency. A small number of concurrent workers (for example, 4–16) per export job is usually enough to saturate ingest capacity without overwhelming downstream processing. Increase only if you observe headroom; back off on 5xx.
  • Stable ordering does not matter. Items can be sent in any order. The upsert keys on (site_id, item_id), not on send order.
  • Plan for resumability. Track which item_ids have been successfully acknowledged so a partially-completed run can resume without re-sending the whole catalog. Resending is safe (see Idempotency below), but skipping already-sent items shortens recovery time.
  • Backfills should run before any user-facing recommendations are exposed. See the Pre-launch phase of the Onboarding Checklist.

Idempotency

A publish payload is an upsert keyed on (site_id, item_id). Reposting the same payload simply overwrites the stored item with the same content — there is no duplicate, no version bump visible to the caller, and no downstream double-counting. This means:

  • Retrying a request after a transport error is safe.
  • Re-running a partial bulk load against the same export is safe.
  • Re-running the bulk load after editing your export (e.g. adding a missing categories value) is the supported way to correct catalog metadata.

This guarantee only holds while the item_id for a given content piece stays the same between runs — re-running with a different item_id orphans the original record. And do not send action: "delete" during a bulk load of published content. For the full republish and soft-delete semantics, see the Item ID Guidance.

Verification

The endpoint returns 202 Accepted as soon as the payload is durably committed to the ingestion pipeline — not when the item is fully visible to the model. To verify that a load completed:

  • Track per-item acknowledgements client-side. A successful 202 per item, with a per-item count matching the export’s row count, is the canonical “done” signal for the ingestion step.
  • Spot-check via the Recommendations API. After the load completes, call GET /recommend/v1/recommendations with a test user_id against the same site_id. A non-empty recommendations array confirms items have reached the model. Empty arrays after a known-good bulk load usually indicate a site_id mismatch between ingestion and the read call.
  • Verify a sample of item_ids round-trip cleanly. Pick a handful of IDs from your export, confirm they show up in recommendations responses for diverse synthetic users, and confirm each one resolves in your CMS.
  • Watch error rates in your client. A bulk load with even a small percentage of 422 errors usually points to a systematic mapping bug in the export — investigate before declaring the load done.

Error and retry behavior

The content endpoint follows the shared status contract in Before you start — nothing endpoint-specific to add. One bulk-load nuance: watch your 422 rate across the load. A systematic spike (rather than the odd row) almost always means a mapping bug in your export, so pause and fix the transform before finishing the run.

Request format and examples

The content endpoint accepts a single JSON UpsertContent object per request, sent with Content-Type: application/json. It does not accept CSV, NDJSON, or a JSON array — those are export formats you may have on disk, but the wire format is always one JSON object per POST /collector/v1/content call.

Submit one payload with curl

Save the UpsertContent object as article-001.json:

{
"action": "publish",
"item_id": "ARTICLE-001",
"site_id": "acme-news",
"type": "article",
"timestamp": "2026-04-08T09:00:00+00:00",
"title": "Election results, day three",
"categories": ["Politics"],
"tags": ["election-2026", "results"],
"author": "Jane Doe",
"is_premium": false,
"metadata": {}
}

POST it with curl, reading the body from the file via -d @:

Terminal window
curl -sS -X POST \
"https://{org}-config-prod.api.arc-cdn.net/collector/v1/content" \
-H "Content-Type: application/json" \
-H "X-API-Key: $COLLECTOR_API_TOKEN" \
-d @article-001.json

A successful submission returns:

HTTP/1.1 202 Accepted
Content-Length: 0

For one-off smoke tests, pass the body inline with -d '{...}' instead of -d @file.

Driving the load from common export formats

Whatever shape your export arrives in, the client-side loop is the same: parse one record, materialise one UpsertContent JSON object, POST it, record the 202. The only thing that changes per format is the parser at the head of the pipeline:

  • NDJSON (one JSON object per line). Stream the file line by line; each line is already a request body. No transformation is needed if the producer writes UpsertContent-shaped objects directly.
  • JSON array. Stream-parse with jq -c '.[]' to emit one object per line, then treat the output as NDJSON.
  • CSV. Map each row to an UpsertContent object. List-valued fields (categories, tags) need a delimiter convention in the CSV — typically ; or | — that the client splits into a JSON array before sending.
  • Database dump. Page through the source table and project each row into an UpsertContent object the same way.

Repeat per record, applying the chunking and concurrency guidance above. Track the item_ids that returned 202 so you can resume without re-sending the whole export if the job is interrupted.

When not to use the content endpoint

  • Ongoing publishes, updates, and deletes. Once the back catalog is loaded, wire the live CMS webhook (or the IFX Recipe for Arc XP CMS customers) so changes propagate within seconds. Re-running the bulk load is not a substitute for the live sync.
  • User-behavior events. Use POST /collector/v1/events or POST /collector/v1/events/batch — see Forwarding events from a CDP or server-side source below and the Compass Batch API guide.
  • Tombstoning items deleted in the source CMS. Bulk load is insert-or- upsert only. To remove items, send action: "delete" payloads explicitly.

Forwarding events from a CDP or server-side source

This half explains how to translate your CDP, analytics, or homegrown event source into the Compass event envelope and POST it to the Collector. Auth, the events endpoints, and retry behavior are the shared rules in Before you start.

What Compass needs from your CDP

Compass trains its recommendation models on user–item interactions. Whatever your source emits, the translation you write needs to answer four questions for every event:

NeedsMeaningExamples
WhoA stable identifier for the user (user_id)logged-in ID, cookie ID, CDP profile ID
WhatThe content item interacted with (item_id)your CMS’s internal record ID (e.g. WordPress post_id, Arc XP _id)
HowThe kind of interaction (event_type)page view, click, save
WhenA timezone-aware timestamp (timestamp)2026-03-27T14:30:00+00:00

Event taxonomy

Compass accepts a fixed set of event_type values — the Supported Event Types table in the API Developer Guide is the authoritative list. On the CDP side your job is to map your source’s native event names onto those values and drop events that don’t map (ad-stack telemetry, consent pings, etc.) rather than forcing them in.

Identity model

Map your source’s stable visitor identifier onto user_id — a CDP-assigned profile ID or a persistent cookie ID works, a per-page-load random value does not. If a visitor logs in mid-session and your source switches from a cookie ID to an account ID, those look like two different users, so decide an identity-resolution strategy on your side. The full anonymization and stability rules are in the User ID Guidance.

item_id conventions

Use your CMS’s internal record ID (e.g. WordPress post_id, Arc XP _id) as item_id, applied consistently, and match the value the content catalog uses for the same item — see A stable item_id underpins everything and the Item ID Guidance.

Batching guidance

Batch for throughput, but don’t sit on events: behavioral recency matters for recommendations. A few seconds to a few minutes of buffering is reasonable for a forwarder; multi-hour batch windows materially degrade recommendation freshness (see the analytics-only section for the worst case). Each batch request is capped at 100 events — see the Batch API for the request format and limits.

Timestamp conventions

The forwarder-specific rule worth repeating: send the time the interaction actually occurred, not the time your forwarder processed it — stale timestamps skew the recency signal that drives recommendations. Prefer UTC; a wrong offset shifts the event in time just as badly. (The base format rule — timezone-aware ISO 8601, naive timestamps rejected — is in the envelope field reference.)

The event envelope

Every event you send is one event-envelope JSON object. The full field reference — types, required vs. optional, and the 422 validation rules — lives in the API Developer Guide (single events) and the Batch API (batched events). This section covers only what you need to produce the shape from a CDP or analytics source.

A minimal valid envelope is the four required fields — user_id, item_id, event_type, and a timezone-aware timestamp:

{
"user_id": "abc-123",
"item_id": "ZSGXFR2KNFCMPN3VHPWQR3BGCE",
"event_type": "page_view",
"timestamp": "2026-03-27T14:30:00+00:00"
}

event_value (a 0.0–1.0 ratio such as scroll depth — not a raw count or duration) and session_id are the optional fields you’re most likely to add:

{
"user_id": "abc-123",
"item_id": "ZSGXFR2KNFCMPN3VHPWQR3BGCE",
"event_type": "deepest_scroll",
"timestamp": "2026-03-27T14:32:00+00:00",
"event_value": 0.75,
"session_id": "sess-a1b2c3"
}

Don’t invent fields outside the documented schema, and leave schema_version unset — any value other than 1 is rejected by design, so old shapes can’t silently degrade analytics.

event_id and de-duplication

If you omit event_id, the Collector fills it deterministically from a hash of the event’s natural identity (user_id, item_id, event_type, timestamp, session_id, and any attribution ID). Under at-least-once delivery — a forwarder retry, a redelivered S3 object — two copies of the same logical event hash to the same event_id, so downstream de-dup treats them as one. The hash only collapses retries whose fields are identical: if you mutate any of those fields between retries (most commonly by stamping a fresh “now” instead of the original interaction time), Compass sees a different event. We recommend omitting event_id and letting Compass derive it, unless your source emits its own stable per-event ID you specifically want preserved end to end (then set event_id to that value).

Where to send it

Send events to POST /collector/v1/events (one envelope) or POST /collector/v1/events/batch (a JSON array of up to 100 envelopes), authenticated with your Collector key in X-API-Key — see Endpoints and Authentication.

Connecting BlueConic

BlueConic emits events via a customer-templated HTTP webhook destination (Mustache). You configure the webhook to render each outbound event directly as an event envelope, so no separate forwarder is needed.

1. Use the Compass BlueConic Mustache template

Download the BlueConic template bundle — a ready-to-paste blueconic.mustache with field-mapping notes in its README.md.

2. Configure the BlueConic webhook destination

In BlueConic (Connections → add a Webhook/HTTP connection):

  1. Destination URL: https://{org}-config-prod.api.arc-cdn.net/collector/v1/events.
  2. Method: POST, Content-Type: application/json.
  3. Headers: add X-API-Key: <your-api-key>.
  4. Payload: paste the Mustache template; map your BlueConic profile/event properties onto the envelope fields (the user identifier → user_id, the content URL/ID → item_id, the BlueConic event type → an event_type).
  5. Timestamp: ensure the rendered timestamp is timezone-aware ISO 8601.
  6. Set the connection’s trigger/goal so it fires on the behavioral events you want Compass to learn from.

BlueConic posts one event per webhook fire to the single-event endpoint; you don’t manage batching yourself. Then verify events are arriving.

Connecting Amplitude / Adobe RT-CDP / Tealium / other HTTP CDP

If your CDP can call an outbound HTTP destination but can’t render the event-envelope shape itself (unlike BlueConic’s Mustache), run a small forwarder: it receives or polls your CDP’s events, translates each to the event envelope, and POSTs batches to Compass. The snippet below is a complete, self-contained forwarder core — fill in translate() with your vendor’s field mapping.

This mirrors the batching/retry conventions of the reference S3 forwarder template so the two stay consistent.

"""Minimal Content Recommendations HTTP forwarder. Translate your CDP records, then POST in batches of <=100."""
import json
import time
import urllib.error
import urllib.request
COLLECTOR_URL = "https://{org}-config-prod.api.arc-cdn.net"
BATCH_PATH = "/collector/v1/events/batch"
HEADERS = {
"X-API-Key": "<your-api-key>", # load from a secret store; never hard-code
"Content-Type": "application/json",
}
MAX_BATCH = 100
MAX_ATTEMPTS = 3
def translate(record: dict) -> dict | None:
"""Map ONE of your CDP's records to the event envelope. Return None to drop it.
Replace the right-hand sides with your vendor's field paths. Drop events
that don't map to the Content Recommendations taxonomy (ad telemetry, consent pings, etc.).
"""
event_type = {"pageview": "page_view", "click": "click"}.get(record.get("type"))
if event_type is None:
return None
return {
"user_id": record["user_id"], # stable per-visitor identifier
"item_id": record["content_id"], # your CMS's internal record ID (stable primary key)
"event_type": event_type,
"timestamp": record["time"], # tz-aware ISO 8601, e.g. 2026-03-27T14:30:00Z
# Optional: "session_id", "event_value" (0.0-1.0). Omit "event_id" to let the collector derive it.
}
def _post_batch(batch: list[dict]) -> None:
body = json.dumps(batch).encode("utf-8")
req = urllib.request.Request(COLLECTOR_URL + BATCH_PATH, data=body, headers=HEADERS, method="POST")
for attempt in range(MAX_ATTEMPTS):
try:
with urllib.request.urlopen(req, timeout=10) as resp: # noqa: S310
resp.read()
return
except urllib.error.HTTPError as exc:
if not (exc.code == 429 or 500 <= exc.code < 600):
raise # 4xx (e.g. 422 bad envelope) is permanent — fix the mapping
except OSError:
pass # transient network/timeout error — retry
time.sleep(2**attempt) # backoff: 1s, 2s, 4s
raise RuntimeError("batch failed after retries")
def forward(records: list[dict]) -> None:
envelopes = [e for r in records if (e := translate(r)) is not None]
for i in range(0, len(envelopes), MAX_BATCH):
_post_batch(envelopes[i : i + MAX_BATCH])

Key points: batch to at most 100, retry only on 5xx / 429 / network errors (a 422 means a bad envelope — fix translate(), don’t retry), and set the X-API-Key auth header on every request (load it from a secret store). This is the shared error and retry behavior applied in code.

Per-vendor notes

  • Amplitude — Use the Event Streaming / forwarding destination to your forwarder, or pull via the Export API. Amplitude’s user_id / device_id split maps to user_id: prefer user_id when present, fall back to device_id. Map event_type names onto the Compass taxonomy in translate().
  • Adobe Real-Time CDP — Use a streaming destination or the Edge Network to reach your forwarder. Resolve an XDM identity into a stable user_id; map XDM web.webPageDetails/media.* events onto the taxonomy.
  • Tealium — Use an EventStream connector (or a webhook destination from Tealium iQ) to your forwarder. The data-layer variable you use for identity becomes user_id; the page/content URL becomes item_id.
  • Any other HTTP CDP — Same shape: get the events to your forwarder, write translate(), POST batches.

Connecting an S3-transport CDP (Permutive, ActionIQ S3)

Some CDPs don’t offer an HTTP webhook destination — they stream activations to an S3 bucket as gzipped NDJSON (Permutive today; ActionIQ S3 activations potentially). Compass does not ingest from object storage. Instead, run a small forwarder Lambda in your own AWS account that reads your bucket and POSTs batches to Compass.

Compass ships a runnable, forkable reference template for exactly this — the S3 forwarder Lambda (s3-forwarder). Download the S3 forwarder template bundle to fork into your own account.

It carries all the S3-side plumbing — s3:ObjectCreated trigger, gunzip + NDJSON parse, batching to 100, Secrets-Manager-backed auth, exponential-backoff retry, and an SQS dead-letter queue. You fork it into your account and write a translate() adapter for your CDP’s record shape (the template includes Permutive-specific field mappings as a worked example). Its bundled README covers deploy steps.

Connecting an analytics-only stack (GA4, Sophi, other)

Some sites collect behavioral events in an analytics tool, not a CDP — most commonly Google Analytics 4 (GA4) or Sophi. These tools don’t sit between your site and a destination the way a CDP does, so there’s no webhook to point at Compass. The integration path is a warehouse export: read the events your analytics tool lands in its data warehouse and forward them to Compass from your own cloud.

GA4 from BigQuery (warehouse export)

Read GA4’s BigQuery export and forward to Compass from your own GCP project. The pattern — a scheduled query / Cloud Function over the daily events_YYYYMMDD table — inlines below (it’s small enough not to warrant a separate template directory).

"""GA4 BigQuery -> Content Recommendations forwarder. Run as a scheduled Cloud Function in YOUR GCP project.
Reads the prior day's events_YYYYMMDD export table, maps GA4 rows to the event envelope,
and POSTs to /events/batch (<=100 per request). The daily table is next-day batch; for
fresher data enable streaming export and read events_intraday_YYYYMMDD instead.
"""
import datetime
import json
import time
import urllib.error
import urllib.request
from google.cloud import bigquery # type: ignore[import-untyped]
COLLECTOR_URL = "https://{org}-config-prod.api.arc-cdn.net"
HEADERS = {
"X-API-Key": "<your-api-key>", # load from Secret Manager; never hard-code
"Content-Type": "application/json",
}
GA4_DATASET = "analytics_000000000" # your GA4 BigQuery export dataset
MAX_BATCH = 100
# Map GA4 event_name -> Content Recommendations event_type. Drop everything not listed.
# Map your org-defined GA4 events (saves, scroll depth, engaged reads) onto
# article_save / deepest_scroll / engaged_read as appropriate.
_EVENT_TYPES = {
"page_view": "page_view",
"select_content": "click",
"search": "search",
"share": "share",
}
def _string_param(row: dict, key: str) -> str | None:
"""Pull a string event-param value out of GA4's repeated event_params array."""
for p in row.get("event_params", []):
if p["key"] == key:
return p["value"].get("string_value")
return None
def _to_envelope(row: dict) -> dict | None:
event_type = _EVENT_TYPES.get(row["event_name"])
item_id = _string_param(row, "content_id") # stable CMS record-ID param; do not fall back to the URL
if event_type is None or not item_id or not row.get("user_pseudo_id"):
return None
# GA4 event_timestamp is epoch microseconds, UTC.
ts = datetime.datetime.fromtimestamp(row["event_timestamp"] / 1_000_000, tz=datetime.UTC)
return {
"user_id": row.get("user_id") or row["user_pseudo_id"], # logged-in id, else pseudo id
"item_id": item_id,
"event_type": event_type,
"timestamp": ts.isoformat(), # tz-aware ISO 8601
}
def _post_batch(batch: list[dict]) -> None:
body = json.dumps(batch).encode("utf-8")
req = urllib.request.Request(
COLLECTOR_URL + "/collector/v1/events/batch", data=body, headers=HEADERS, method="POST"
)
for attempt in range(3):
try:
with urllib.request.urlopen(req, timeout=10) as resp: # noqa: S310
resp.read()
return
except urllib.error.HTTPError as exc:
if not (exc.code == 429 or 500 <= exc.code < 600):
raise # permanent (e.g. 422 bad envelope)
except OSError:
pass # transient network/timeout error — retry
time.sleep(2**attempt)
raise RuntimeError("batch failed after retries")
def run(yyyymmdd: str) -> None:
client = bigquery.Client()
table = f"{GA4_DATASET}.events_{yyyymmdd}"
rows = client.query(f"SELECT * FROM `{table}`").result() # noqa: S608 - dataset is your own config
batch: list[dict] = []
for row in rows:
envelope = _to_envelope(dict(row))
if envelope is None:
continue
batch.append(envelope)
if len(batch) == MAX_BATCH:
_post_batch(batch)
batch = []
if batch:
_post_batch(batch)

Deploy as a Cloud Function triggered by Cloud Scheduler (e.g. daily after the export lands). You operate this in your own GCP project — Compass doesn’t run it for you.

Sophi integration note

Sophi collects through its own SDK / hosted pipeline. Unverified (no live vendor access at time of writing): whether Sophi exposes an outbound real-time webhook or a documented streaming/partner export API is not confirmed. If your Sophi contract includes a webhook or export, you can forward from it using the same HTTP forwarder snippet (or the S3 forwarder if the export is to S3). Confirm Sophi’s actual export capabilities with your Sophi account team before building a server-side forwarder.

Connecting a homegrown event source

If you emit events from your own application code or a homegrown pipeline, you’re already in control of the shape — just produce the event envelope and POST it. Use the inline HTTP forwarder snippet as your starting point: it shows the batching, retry, and header conventions Compass expects. The only part specific to you is translate() — and if your code can emit the envelope directly, you can skip translation and POST envelopes straight to /events/batch.

Verifying events are arriving

After wiring up any path above:

  1. Send a known test event. POST a single, hand-built envelope to /collector/v1/events with a recognizable user_id (e.g. qa-smoke-001) and a current timestamp.

    Terminal window
    curl -i -X POST \
    "https://{org}-config-prod.api.arc-cdn.net/collector/v1/events" \
    -H "X-API-Key: <your-api-key>" \
    -H "Content-Type: application/json" \
    -d '{"user_id":"qa-smoke-001","item_id":"test-item-1","event_type":"page_view","timestamp":"2026-03-27T14:30:00+00:00"}'
  2. Confirm the response. A 2xx (typically 202/200) means the envelope was accepted. A 422 is a validation failure — the response body names the offending field; fix it (most often a naive timestamp or an unknown event_type).

  3. Confirm auth. A 401/403 means a wrong X-API-Key or the wrong host. Re-check your key and base URL.

  4. Confirm volume. Once live, your forwarder/webhook logs (or the S3 forwarder’s CloudWatch processed_object lines, or the GA4 function logs) should show steady accepted batches with few or no 422s.

  5. What success looks like: a steady stream of 2xx responses, validation-failure rate near zero, and recommendation quality improving as interaction history accumulates. (New tenants go through a cold-start period before history is rich enough — ask your Arc contact about expected ramp.)

Debugging quick reference:

SymptomLikely causeFix
422 on every eventBad envelope (naive timestamp, unknown event type)Read the response body; fix translate().
401 / 403Wrong X-API-Key, or wrong hostRe-verify your key and the base URL with your Arc contact.
2xx but no recommendation liftuser_id not stable, or events all droppedConfirm translate() isn’t returning None for everything.
Recommendations feel staleBatch window too long / warehouse exportShorten buffering; enable streaming export instead of next-day batch.

Common pitfalls

  • Anonymous vs identified users. A login transition splits one reader into two IDs unless you stitch identity on your side — see the User ID Guidance for the workable patterns.
  • Event de-dup behavior. Covered under event_id and de-duplication — retries must resend identical field values, including the original timestamp.
  • Timestamp timezones. Covered under Timestamp conventions — timezone-aware ISO 8601, original interaction time, prefer UTC.
  • Batch size limits. /events/batch accepts at most 100 events per request. Slice larger sets. A batch of 101 is rejected — not silently truncated.
  • event_value and schema_version. Both covered in the event envelope — event_value is a 0.0–1.0 ratio, and schema_version stays unset.
  • Don’t force-fit events. Drop events that don’t map to the Compass taxonomy rather than inventing an event_type — and never invent envelope fields outside the contract.

See also