Skip to content

Compass ID Guidance

Audience: Tenant integrators wiring a CMS to the Compass ingestion API, and anyone sending behavioral events to — or reading recommendations from — Compass.

Every interaction the recommendation model trains on is a pair: an item_id (which piece of content) and a user_id (which reader). The model learns “this reader interacted with this item,” and both halves of that pair must stay stable and consistent for that learning to accumulate. Change either identifier and the interaction history keyed on the old value is orphaned — the model cannot tell that the new value refers to the same content or the same person, and continuity restarts from zero.

Neither identifier may ever contain PII: both item_id and user_id surface in operational logs and internal data exports, and Arc XP performs no PII scrubbing on ingestion. Hash or pseudonymize anything derived from personal data before sending.

This page is the single reference for both identifiers. Other Compass guides link here instead of repeating these rules.

Item ID

The choice of item identifier is the single most consequential decision a tenant makes during onboarding. The model’s learning is keyed on the item ID you supplied at ingest time, so a reissued ID orphans everything recorded against the old one (see above) and restarts that item’s training continuity from zero.

Get this right at onboarding and you will rarely think about it again. Get it wrong, and every CMS edit that reissues an ID quietly degrades your model.

TL;DR

  • Pick one ID source per content piece, on Day 1, and never change it. Treat it as a primary key, not a label.
  • Prefer the CMS’s internal record ID over the canonical URL unless your CMS guarantees URL stability across moves, renames, and re-categorization.
  • Use the API constraint ceiling as a hard cap, not a target. API requests reject IDs over 256 characters; aim for 64 characters or fewer to keep logs and dashboards readable.
  • Never reuse a deleted item’s ID. If content needs to return, mint a new ID. The recommendation engine cannot tell “the old one came back” from “a different item appeared at the same key” — it will fold both histories together.

What “item ID” means in this system

When you POST to /collector/v1/content, the value you send as item_id flows through to the recommendation engine as the canonical key for that piece of content. The same identifier is what comes back as recommendations[].item_id when you call the recommendations endpoint, and it is the join key used to attach the display-ready card payload (title, author, URL, thumbnail, etc.) to each recommendation.

Three places see the same value:

  1. Your CMS — the source of truth for the content piece.
  2. Our ingestion API (item_id field on content events).
  3. The recommendation engine (the catalog key against which user interactions are recorded, and the key by which each recommendation’s card is looked up from the stored Item document).

For recommendations to mean what you expect, all three must use exactly the same string for the same piece of content, every time. A mismatch does not just orphan interaction history — it also returns card: null for the recommendation, because the lookup against the Item document misses.

Stability requirements

Immutability

The ID must remain identical for the entire lifetime of the content piece. Once you have published an item with item_id = "abc123" and a single user interaction has been recorded against it, you should treat that string as permanent. Republishing the same article with a different ID is treated as a new item with zero history.

This is a customer-side responsibility. The platform does not detect “this new item is really the same as that old one.” If your CMS workflow can mint new IDs on edits, on locale changes, on category moves, or on re-publishes, that workflow is incompatible with stable recommendations as-is and needs to be reviewed.

IDs are compared byte-for-byte. abc123 and ABC123 are two different items, as are abc-123 and abc_123. Pick a casing and punctuation convention up front and apply it consistently — silently changing case on a republish has the same effect as changing the ID outright.

Character set

The platform does not enforce a character-set restriction on item_id. Any UTF-8 string of an allowed length is accepted by the API. You are responsible for the format you choose. We recommend:

  • ASCII alphanumerics, plus -, _, ., /, :.
  • Avoid whitespace, control characters, and characters that need URL-encoding in path segments — IDs surface in logs, dashboards, and debug tooling.
  • Avoid characters that may collide with downstream serialization (newlines, tabs, NUL bytes).

Length

The API enforces a length bound of 1–256 characters at request time. Requests with an item_id outside that range are rejected with a 422.

Treat that bound as authoritative: 1–256 characters, target ≤ 64. Longer IDs work but make operational tooling harder to read, and risk truncation in third-party exports.

Choosing your ID source: use the CMS internal ID

Use the primary key your CMS assigns to the record (e.g. WordPress post_id, Drupal nid; for Arc XP CMS content, the Arc ID is used automatically — see the note above). CMSes treat their PKs as immutable, so this ID survives the edits that most often break recommendation continuity: slug rewrites, SEO URL changes, re-categorization, and moves between sections.

Do not use the canonical URL or its slug as item_id. URLs are content, not keys — editorial workflows routinely rewrite them (“better SEO,” “fix typo,” re-categorization from /news/foo to /politics/foo), and each rewrite orphans every interaction recorded against the old URL.

Watch-outs when using the CMS PK:

  • Tied to one CMS. A platform migration may reset PKs and orphan all history. If you ever change CMSes, plan an ID-mapping step at migration time so the old PKs continue to resolve to the same items under the new system.

Republish behavior

The ingestion path is idempotent on (tenant_id, site_id, item_id). Sending a content event for an existing item updates the stored record in place; sending an event for a new combination creates a new item. tenant_id is supplied per request out of band (resolved from your credentials / base URL), not carried in the item payload — so from a producer’s point of view the payload only needs site_id and item_id.

Same item_id on republish:

  • Item record is updated (title, body, categories, etc. refresh).
  • The recommendation engine’s catalog entry is upserted under the same key.
  • Interaction history is preserved. Existing model knowledge carries forward.
  • The next recommendations response surfaces the refreshed fields on the card payload for that item.

New item_id on republish (do not do this):

  • A new item appears in your catalog.
  • The recommendation engine learns about a brand-new item with zero interactions.
  • The old item still exists with all its history but is no longer being surfaced by your CMS — it sits in the catalog as a dead entry.
  • Recommendations regress for the duration of training catch-up.
  • If the old ID is still recommended (e.g. via an editorial pin) before the new document syncs, its card will be null in the response and front-ends will skip rendering it.

If your CMS’s republish workflow assigns a new ID — whether due to versioning, “unpublish then republish” actions, or import from another source — that workflow needs to either preserve the original ID or be replaced before the content reaches our ingestion endpoint.

Delete semantics and the takedown pattern

Why delete-and-recreate is destructive

The recommendation engine learns from interaction sequences over time. The training signal that makes recommendations useful — “users who watched A, B, and C went on to read D” — is keyed on the exact item-ID strings recorded against past interactions.

Deleting an item and recreating it under the same ID after a delay is not equivalent to leaving it untouched. The interaction history still references the original ID, but in the meantime the model has been retrained on a catalog without it; the new item appears as a fresh, untrained entry that will not surface in recommendations until enough new interactions accumulate.

Worse, delete-and-recreate with a different ID orphans every interaction ever recorded against the original ID. There is no retroactive “rebind.”

For voluntary takedowns (article retired, content out of date):

  1. Soft-delete the item by sending a delete action through the content ingestion endpoint. The catalog retains the record and the training history is preserved.
  2. The recommendations endpoint will exclude soft-deleted items from results.
  3. Do not reuse the ID if the content is later restored under a new editorial workflow — mint a new ID for the new item.

API surface, in one place

FieldWhereConstraint
item_idPOST /collector/v1/content request1–256 chars; UTF-8; no charset filter
recommendations[].item_idGET /recommend/v1/recommendations responsestring echoed verbatim from what was ingested
recommendations[].cardGET /recommend/v1/recommendations responsenested display payload looked up by (tenant_id, site_id, item_id); null on lookup miss

User ID

user_id is the key the model learns on for each reader. Every event you send and every recommendations request you make carries one, and personalization quality is a direct function of two properties you control: the ID is anonymized, and the ID is stable.

TL;DR

  • Never send PII. user_id must be an anonymized token — not an email, name, phone number, or any directly identifying value. Hash or pseudonymize before sending. Arc XP does not sanitize user IDs on ingestion.
  • Same reader, same ID, every time. The same person must resolve to the same user_id across sessions — and across devices and collection paths where possible. Anonymous IDs are fine as long as they are stable.
  • A per-page-load random value is the worst case. It makes every event look like a brand-new user, so the model can never accumulate history.
  • Use the same raw value when sending events and when fetching recommendations. The write path and the read path pseudonymize user_id identically — if the two paths send different raw values, personalization silently breaks.

Anonymization

user_id may be an authenticated account ID or an ephemeral ID for anonymous readers, but it must be anonymized before it reaches Arc XP. Hash or pseudonymize on your side; the platform stores what you send.

This is a hard requirement, not a recommendation — for the reasons given in the intro: IDs surface in logs and exports, and Arc XP performs no PII scrubbing on ingestion. The item-ID format guidance above applies to user_id too: keep it to readable ASCII, free of whitespace and PII, so it stays legible wherever it surfaces.

Stability

The model builds a per-reader interaction history keyed on user_id. Every place an ID can silently change is a place that history fragments:

  • Across sessions. A returning visitor must produce the same ID as their previous visit. A persistent first-party cookie or a CDP-assigned profile ID works; a session-scoped or per-page-load ID does not.
  • Across devices, where you can manage it. A stable, host-managed ID (for example, a hashed logged-in ID) is the only way one reader’s phone and laptop share a single history.
  • Across collection paths. If you run the Compass Web SDK for live web traffic and a forwarder for backfill or non-web surfaces, both must emit the same user_id for the same reader, or the model treats the two streams as unrelated people.
  • Between backfill and live traffic. Historical events replayed during onboarding must use the same anonymization scheme your production traffic will use, so pre-launch history attaches to the readers who show up on day one.

One value across write and read

The user_id you put on events must be the same raw value your recommendations requests use. Both paths pseudonymize identically, so a mismatch doesn’t error — it just returns recommendations for a reader with no history.

If you use the Compass Web SDK, read the value back with getUserId() and forward it verbatim to the read path. See Compass Web SDK: Consent & Identity.

Login transitions and identity stitching

When a visitor logs in, many event sources switch from a cookie or device ID to an account ID. Compass performs no identity stitching: the pre-login history stays attached to the old ID, and the account ID starts from zero.

Decide a strategy on your side and apply it consistently. The two workable patterns:

  • Keep emitting the pre-login ID as user_id even after login, or
  • Maintain your own mapping and always resolve a visitor to one canonical ID before sending.

Whichever you pick, apply it on both the event path and the recommendations read path.

What happens when you get it wrong

Unstable or inconsistent IDs don’t produce errors — they produce quietly bad results:

  • The model can’t accumulate history per reader, so results stay generic and popularity-based long after launch.
  • The same reader sees different personalization on different surfaces or devices.
  • Backfilled history never attaches to live readers, wasting the seeding work.

If recommendations look low-relevance, user_id consistency is one of the first things to audit — see the Troubleshooting section of the API Developer Guide.

New and anonymous readers

A reader with no history is expected, not a failure: Compass returns popularity-based results for them automatically, and results personalize progressively as their events accumulate. You don’t need special handling — you need their ID to stay stable so the accumulation happens.

See also