Compass ID Guidance
Audience: Tenant integrators wiring a CMS to the Compass ingestion API, and anyone sending behavioral events to — or reading recommendations from — Compass.
Every interaction the recommendation model trains on is a pair: an item_id (which piece of content) and a user_id (which reader). The model learns “this reader interacted with this item,” and both halves of that pair must stay stable and consistent for that learning to accumulate. Change either identifier and the interaction history keyed on the old value is orphaned — the model cannot tell that the new value refers to the same content or the same person, and continuity restarts from zero.
Neither identifier may ever contain PII: both item_id and user_id surface in operational logs and internal data exports, and Arc XP performs no PII scrubbing on ingestion. Hash or pseudonymize anything derived from personal data before sending.
This page is the single reference for both identifiers. Other Compass guides link here instead of repeating these rules.
Item ID
The choice of item identifier is the single most consequential decision a tenant makes during onboarding. The model’s learning is keyed on the item ID you supplied at ingest time, so a reissued ID orphans everything recorded against the old one (see above) and restarts that item’s training continuity from zero.
Get this right at onboarding and you will rarely think about it again. Get it wrong, and every CMS edit that reissues an ID quietly degrades your model.
TL;DR
- Pick one ID source per content piece, on Day 1, and never change it. Treat it as a primary key, not a label.
- Prefer the CMS’s internal record ID over the canonical URL unless your CMS guarantees URL stability across moves, renames, and re-categorization.
- Use the API constraint ceiling as a hard cap, not a target. API requests reject IDs over 256 characters; aim for 64 characters or fewer to keep logs and dashboards readable.
- Never reuse a deleted item’s ID. If content needs to return, mint a new ID. The recommendation engine cannot tell “the old one came back” from “a different item appeared at the same key” — it will fold both histories together.
What “item ID” means in this system
When you POST to /collector/v1/content, the value you send as item_id flows through to the recommendation engine as the canonical key for that piece of content. The same identifier is what comes back as recommendations[].item_id when you call the recommendations endpoint, and it is the join key used to attach the display-ready card payload (title, author, URL, thumbnail, etc.) to each recommendation.
Three places see the same value:
- Your CMS — the source of truth for the content piece.
- Our ingestion API (
item_idfield on content events). - The recommendation engine (the catalog key against which user interactions are recorded, and the key by which each recommendation’s
cardis looked up from the stored Item document).
For recommendations to mean what you expect, all three must use exactly the same string for the same piece of content, every time. A mismatch does not just orphan interaction history — it also returns card: null for the recommendation, because the lookup against the Item document misses.
Stability requirements
Immutability
The ID must remain identical for the entire lifetime of the content piece. Once you have published an item with item_id = "abc123" and a single user interaction has been recorded against it, you should treat that string as permanent. Republishing the same article with a different ID is treated as a new item with zero history.
This is a customer-side responsibility. The platform does not detect “this new item is really the same as that old one.” If your CMS workflow can mint new IDs on edits, on locale changes, on category moves, or on re-publishes, that workflow is incompatible with stable recommendations as-is and needs to be reviewed.
IDs are compared byte-for-byte. abc123 and ABC123 are two different items, as are abc-123 and abc_123. Pick a casing and punctuation convention up front and apply it consistently — silently changing case on a republish has the same effect as changing the ID outright.
Character set
The platform does not enforce a character-set restriction on item_id. Any UTF-8 string of an allowed length is accepted by the API. You are responsible for the format you choose. We recommend:
- ASCII alphanumerics, plus
-,_,.,/,:. - Avoid whitespace, control characters, and characters that need URL-encoding in path segments — IDs surface in logs, dashboards, and debug tooling.
- Avoid characters that may collide with downstream serialization (newlines, tabs, NUL bytes).
Length
The API enforces a length bound of 1–256 characters at request time. Requests with an item_id outside that range are rejected with a 422.
Treat that bound as authoritative: 1–256 characters, target ≤ 64. Longer IDs work but make operational tooling harder to read, and risk truncation in third-party exports.
Choosing your ID source: use the CMS internal ID
Use the primary key your CMS assigns to the record (e.g. WordPress post_id, Drupal nid; for Arc XP CMS content, the Arc ID is used automatically — see the note above). CMSes treat their PKs as immutable, so this ID survives the edits that most often break recommendation continuity: slug rewrites, SEO URL changes, re-categorization, and moves between sections.
Do not use the canonical URL or its slug as item_id. URLs are content, not keys — editorial workflows routinely rewrite them (“better SEO,” “fix typo,” re-categorization from /news/foo to /politics/foo), and each rewrite orphans every interaction recorded against the old URL.
Watch-outs when using the CMS PK:
- Tied to one CMS. A platform migration may reset PKs and orphan all history. If you ever change CMSes, plan an ID-mapping step at migration time so the old PKs continue to resolve to the same items under the new system.
Republish behavior
The ingestion path is idempotent on (tenant_id, site_id, item_id). Sending a content event for an existing item updates the stored record in place; sending an event for a new combination creates a new item. tenant_id is supplied per request out of band (resolved from your credentials / base URL), not carried in the item payload — so from a producer’s point of view the payload only needs site_id and item_id.
Same item_id on republish:
- Item record is updated (title, body, categories, etc. refresh).
- The recommendation engine’s catalog entry is upserted under the same key.
- Interaction history is preserved. Existing model knowledge carries forward.
- The next recommendations response surfaces the refreshed fields on the
cardpayload for that item.
New item_id on republish (do not do this):
- A new item appears in your catalog.
- The recommendation engine learns about a brand-new item with zero interactions.
- The old item still exists with all its history but is no longer being surfaced by your CMS — it sits in the catalog as a dead entry.
- Recommendations regress for the duration of training catch-up.
- If the old ID is still recommended (e.g. via an editorial pin) before the new document syncs, its
cardwill benullin the response and front-ends will skip rendering it.
If your CMS’s republish workflow assigns a new ID — whether due to versioning, “unpublish then republish” actions, or import from another source — that workflow needs to either preserve the original ID or be replaced before the content reaches our ingestion endpoint.
Delete semantics and the takedown pattern
Why delete-and-recreate is destructive
The recommendation engine learns from interaction sequences over time. The training signal that makes recommendations useful — “users who watched A, B, and C went on to read D” — is keyed on the exact item-ID strings recorded against past interactions.
Deleting an item and recreating it under the same ID after a delay is not equivalent to leaving it untouched. The interaction history still references the original ID, but in the meantime the model has been retrained on a catalog without it; the new item appears as a fresh, untrained entry that will not surface in recommendations until enough new interactions accumulate.
Worse, delete-and-recreate with a different ID orphans every interaction ever recorded against the original ID. There is no retroactive “rebind.”
Recommended takedown pattern
For voluntary takedowns (article retired, content out of date):
- Soft-delete the item by sending a
deleteaction through the content ingestion endpoint. The catalog retains the record and the training history is preserved. - The recommendations endpoint will exclude soft-deleted items from results.
- Do not reuse the ID if the content is later restored under a new editorial workflow — mint a new ID for the new item.
API surface, in one place
| Field | Where | Constraint |
|---|---|---|
item_id | POST /collector/v1/content request | 1–256 chars; UTF-8; no charset filter |
recommendations[].item_id | GET /recommend/v1/recommendations response | string echoed verbatim from what was ingested |
recommendations[].card | GET /recommend/v1/recommendations response | nested display payload looked up by (tenant_id, site_id, item_id); null on lookup miss |
User ID
user_id is the key the model learns on for each reader. Every event you send and every recommendations request you make carries one, and personalization quality is a direct function of two properties you control: the ID is anonymized, and the ID is stable.
TL;DR
- Never send PII.
user_idmust be an anonymized token — not an email, name, phone number, or any directly identifying value. Hash or pseudonymize before sending. Arc XP does not sanitize user IDs on ingestion. - Same reader, same ID, every time. The same person must resolve to the
same
user_idacross sessions — and across devices and collection paths where possible. Anonymous IDs are fine as long as they are stable. - A per-page-load random value is the worst case. It makes every event look like a brand-new user, so the model can never accumulate history.
- Use the same raw value when sending events and when fetching
recommendations. The write path and the read path pseudonymize
user_ididentically — if the two paths send different raw values, personalization silently breaks.
Anonymization
user_id may be an authenticated account ID or an ephemeral ID for anonymous
readers, but it must be anonymized before it reaches Arc XP. Hash or
pseudonymize on your side; the platform stores what you send.
This is a hard requirement, not a recommendation — for the reasons given in
the intro: IDs surface in logs and exports, and Arc XP performs no PII scrubbing
on ingestion. The item-ID format guidance above applies to user_id too: keep
it to readable ASCII, free of whitespace and PII, so it stays legible wherever
it surfaces.
Stability
The model builds a per-reader interaction history keyed on user_id. Every
place an ID can silently change is a place that history fragments:
- Across sessions. A returning visitor must produce the same ID as their previous visit. A persistent first-party cookie or a CDP-assigned profile ID works; a session-scoped or per-page-load ID does not.
- Across devices, where you can manage it. A stable, host-managed ID (for example, a hashed logged-in ID) is the only way one reader’s phone and laptop share a single history.
- Across collection paths. If you run the Compass Web SDK for live web
traffic and a forwarder for backfill or non-web surfaces, both must emit the
same
user_idfor the same reader, or the model treats the two streams as unrelated people. - Between backfill and live traffic. Historical events replayed during onboarding must use the same anonymization scheme your production traffic will use, so pre-launch history attaches to the readers who show up on day one.
One value across write and read
The user_id you put on events must be the same raw value your
recommendations requests use. Both paths pseudonymize identically, so a
mismatch doesn’t error — it just returns recommendations for a reader with no
history.
If you use the Compass Web SDK, read the value back with getUserId() and
forward it verbatim to the read path. See
Compass Web SDK: Consent & Identity.
Login transitions and identity stitching
When a visitor logs in, many event sources switch from a cookie or device ID to an account ID. Compass performs no identity stitching: the pre-login history stays attached to the old ID, and the account ID starts from zero.
Decide a strategy on your side and apply it consistently. The two workable patterns:
- Keep emitting the pre-login ID as
user_ideven after login, or - Maintain your own mapping and always resolve a visitor to one canonical ID before sending.
Whichever you pick, apply it on both the event path and the recommendations read path.
What happens when you get it wrong
Unstable or inconsistent IDs don’t produce errors — they produce quietly bad results:
- The model can’t accumulate history per reader, so results stay generic and popularity-based long after launch.
- The same reader sees different personalization on different surfaces or devices.
- Backfilled history never attaches to live readers, wasting the seeding work.
If recommendations look low-relevance, user_id consistency is one of the
first things to audit — see the Troubleshooting section of the
API Developer Guide.
New and anonymous readers
A reader with no history is expected, not a failure: Compass returns popularity-based results for them automatically, and results personalize progressively as their events accumulate. You don’t need special handling — you need their ID to stay stable so the accumulation happens.
See also
- Onboarding checklist — Compass Onboarding Checklist for the phase-by-phase rollout steps and where the identifier decisions land.
- OpenAPI / Swagger — see the Compass API Reference for the live schema.
- API overview — Compass API Developer Guide for endpoint-level reference and field details.
- Compass Web SDK: Consent & Identity — Compass Web SDK: Consent & Identity for how the SDK resolves and persists reader identity.