# SegmentStream AI Infrastructure that gives AI agents a marketing measurement brain. Any question about your ads. Any report. Any recommendation. From your AI, not next week's meeting. ## What SegmentStream Does SegmentStream connects marketing data, ad platforms, analytics tools, CRMs, and warehouses to AI agents through the Model Context Protocol. Your AI can run attribution analysis, investigate anomalies, generate reports, recommend budget reallocations, and set up measurement workflows with real SegmentStream tools. ## How It Works - **Connect your data**: Ad platforms, analytics, CRM, data warehouse, and more. - **Add SegmentStream MCP to your AI tool**: One config line in Claude Code, Claude Cowork, Cursor, or any MCP client. - **Start working**: Analysis, optimization, measurement — all through conversation with your AI. ## Analyst-Grade Outputs - **Campaign Performance Report**: Cross-channel performance with attribution-adjusted ROAS, anomaly flags, and week-over-week trends. - **Root Cause Analysis**: Automated investigation of metric anomalies — root cause identification with actionable recommendations. - **Budget Reallocation Plan**: Current vs proposed allocation with marginal ROAS analysis, projected revenue impact, and channel-level reasoning. - **Month-over-Month Analysis**: Automated period comparison with trend charts, statistical significance flags, and channel-level drill-down. ## Built-In Measurement Expertise - **Analysis**: Performance review, Baseline demand analysis, Anomaly detection - **Optimization**: Budget allocation, Bid strategy evaluation, Marginal ROAS analysis - **Setup**: Project setup, Conversion tracking, Data source configuration - **Debugging**: Unattributed investigation, SDK pipeline diagnosis, Data quality audit ## Infrastructure - **Your data stays yours**: All marketing data lives in your own BigQuery warehouse. SegmentStream reads and enriches it — never copies or moves it. Switch away anytime and keep everything. - **Your AI. Any client.**: Works with Claude Code, Claude Cowork, Cursor, or any MCP-compatible client. No proprietary interface to learn — use the AI tools your team already runs. - **The measurement brain**: Cross-channel attribution, budget optimization, incrementality testing, and predictive scoring — exposed as MCP tools your AI can call directly. The methodology that took years to build, accessible in one config line. - **Open protocol. Zero lock-in.**: Built on Model Context Protocol — the open standard for AI tool integration. No vendor lock-in on data, interface, or methodology. Every piece of the stack is replaceable except the measurement intelligence. ## Production Engine Powered by the same measurement engine that manages $100M+ in annual ad spend for enterprise teams. ML attribution models, incrementality testing, and budget optimization — built over years, now accessible through Claude Code, Claude Cowork, Cursor, and any MCP-compatible tool. - **$100M+** ad spend managed - **30+** ad platform integrations - **13** measurement engine capabilities ## Key Pages - [AI Skills](https://segmentstream.com/ai-skills.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) - [Integrations](https://segmentstream.com/integrations.md) - [Pricing](https://segmentstream.com/pricing.md) - [For Agencies](https://segmentstream.com/agency.md) - [About](https://segmentstream.com/about.md) --- # AI Skills Browse the catalogue of curated skills shipped inside the SegmentStream MCP. Copy any prompt into Claude, Codex, or Cursor and get analyst-grade answers grounded in your real data. ## Marketing skills your AI doesn't have out of the box Curated capabilities shipped inside the SegmentStream MCP. Browse by question, copy any prompt into Claude, Codex, or Cursor. ## Set up & connect Create the project, connect analytics, ad platforms, CRM, and warehouse data, then configure reporting logic before your team acts on the numbers. - **Project setup**: `Help me set up my SegmentStream project.` - **Google Analytics connection**: `Help me connect Google Analytics to SegmentStream.` - **Ad platform connections**: `Help me connect Meta Ads and Google Ads to SegmentStream.` - **CRM connection**: `Help me connect my CRM to SegmentStream.` - **Data warehouse connection**: `Help me connect my data warehouse to SegmentStream.` - **Channel grouping setup**: `Help me set up my Channel grouped dimension for paid, organic, referral, direct, and AI Search traffic.` - **Attribution settings**: `Help me review my attribution settings and choose the best model for reporting.` - **Conversion setup**: `Help me create a conversion from my website events.` ## Data quality Check CRM data, cost data, unattributed conversions, and identity coverage before using attribution reports for budget decisions. - **Data quality audit**: `Help me audit my SegmentStream data quality before I rely on attribution reports.` - **CRM data quality**: `Help me check whether my CRM data is connected correctly to attribution reports.` - **Cost data quality**: `Help me check whether my ad cost data is complete and matched correctly.` - **Unattributed conversions**: `Help me understand why some conversions are unattributed and how to improve attribution coverage.` - **Identity graph coverage**: `Help me improve cross-device and cross-browser journeys with SegmentStream Identity Graph.` ## Audit traffic & audience quality Bots, audience leakage, and existing-demand capture distort performance before budget decisions are made. These skills show whether traffic is real, which platforms see your audience, and which channels are mostly reaching demand you already had. - **Bot traffic detection**: `Help me check whether my paid traffic is bots or real people.` - **Audience leakage audit**: `Help me audit which ad platforms can see my website audience and whether that signal leaks value.` - **Self-Reported Attribution classifier**: `Help me review my Self-Reported Attribution classifier.` - **Existing-demand capture**: `Help me find channels that mostly capture existing demand.` ## Compare & decide Two questions every team has to answer before moving budget: which attribution model are we trusting, and which channels are actually carrying their weight. - **Attribution model selection**: `Help me choose the right attribution model for my business.` - **Compare paid channels**: `Help me rank paid channels by ROAS and decide where to scale or cut.` ## Diagnose discrepancies Google Analytics says one thing. SegmentStream says another. First-click and last-click point to different winners. These skills isolate the structural reason — not a guess. - **Google Analytics discrepancy**: `Help me understand why SegmentStream data does not match Google Analytics.` - **First-click vs last-click**: `Help me understand why first-click and last-click attribution differ so much for one channel.` - **Cannibalization investigation**: `Help me check whether retargeting creates new conversions.` ## Test, learn, explore Incrementality results need careful interpretation. Saved reports keep your team from rebuilding the same query every quarter. - **Geo holdout interpretation**: `Help me interpret my geo holdout results.` - **Saved reports**: `Help me run my saved channel performance report.` ## Measurement Engine Explore SegmentStream capabilities, measurement theory, and the cases where attribution needs another method. - **Capabilities overview**: `Help me understand which SegmentStream Measurement Engine capabilities I can use.` - **Predictive attribution**: `Help me understand how SegmentStream predictive attribution works.` - **Self-Reported Reattribution**: `Help me understand how SegmentStream Self-Reported Reattribution works.` - **Post-view effects**: `Help me understand how SegmentStream measures post-view ad effects.` - **Marginal analytics**: `Help me understand marginal ROAS and how SegmentStream uses it for budget decisions.` - **Budget allocation**: `Help me understand how automated budget allocation works in SegmentStream.` - **Predictive scoring**: `Help me understand how SegmentStream predicts lead quality, CRM outcomes, or LTV.` --- # SegmentStream Cloud — White-Label Marketing Analytics Add marketing analytics to your product — or run it for your clients. Attribution, incrementality testing and budget optimization — fully managed on SegmentStream Cloud. Under your brand, fed from any source, priced by what you process. Nothing to deploy. SegmentStream Cloud is the SegmentStream measurement engine offered as an embeddable, white-label engine, managed by SegmentStream. It is currently rolling out to a small group of design partners; the page collects waitlist signups. > **[Interactive: CLI terminal]** > A terminal frame shows the provisioning flow: `npm i -g @segmentstream/cli`, then > `segmentstream workspaces create acme --region eu` (workspace ready), `segmentstream deploy` > (segmentstream.yml deployed — 4 sources, 2 models), and `segmentstream run --workflow refresh` > (refresh complete — identity graph + models updated). Labeled "CLI preview — commands may > change before GA." At a glance: REST + TypeScript SDK · MCP server per workspace · EU & US regions · workspace-per-client isolation · GDPR / DPA. ## Two products, one engine Measuring your own brand's marketing? You're looking for the platform. This page is the white-label version — the same engine, inside your product. - **SegmentStream (the platform)** — for marketing teams and brands. Measures your own marketing. Hosted dashboard and AI agent, run by us. SegmentStream branding. Scoped per engagement, priced after discovery. See [Pricing](https://segmentstream.com/pricing.md). - **SegmentStream Cloud (the white-label engine)** — for SaaS platforms and agencies. Measurement for your customers. White-label module, SDK / API / MCP inside your product. Your branding. Platform fee + DPU / DQU usage. Join the waitlist at https://segmentstream.com/cloud. Same engine underneath — attribution, incrementality testing, budget optimization. The difference is whose brand is on the screen and whose customers are looking at it. ## Two ways to run it Two entry points into Cloud, metered the same way — pick yours by where your events come from. ### Platform — for SaaS platforms You already have a tracker — an adapter maps your existing events; clients never re-tag. - White-label Measurement module inside your app - TypeScript SDK and API to build your own surface - Your brand in front, our engine underneath Example: an ad-platform ships a white-label Measurement tab inside its own admin panel. ### Agency — for agencies & in-house teams No tracker needed — connect GA4 exports or warehouse SQL as event sources. - A workspace per client, the full engine in each - Adobe Analytics next — same adapter pattern - Nothing to deploy on the client's site Example: an agency runs 30 client workspaces off GA4 BigQuery exports. ## Your product. Our engine. Not an iframe — React components and a typed SDK. Build an admin panel in any layout and design system; SegmentStream appears nowhere your customers look. > **[Interactive: composable UI demo]** > A design switcher renders the same engine data as three completely different admin panels: > a sidebar admin ("Acme Ads", channel table + incrementality badge), a KPI dashboard > ("Northwind", top navigation, KPI cards for Cost / Conv. Value / ROAS / Incrementality and a > trend chart), and a minimal custom report ("Your design", big blended-ROAS number rendered by > a 40-line internal script). Layout, navigation and design system differ entirely; the data > and SDK underneath are identical. ## Any source in. Any surface out. Identity resolution, attribution models, incrementality and budget optimization run on managed SegmentStream infrastructure — deployed in the cloud region where your clients live. Sources: - SegmentStream SDK — first-party tracking, or your own tracker via an adapter - GA4 export — BigQuery event export, no re-tagging - Warehouse SQL — any table you can query becomes an event source - Adobe Analytics — next in line, same adapter pattern Surfaces: - White-label module — the full measurement UI inside your product, your brand - REST API & SDK — every report and model output, programmatically - MCP for AI agents — Claude, Cursor or any agent queries the engine directly - Dashboards — hosted workspaces when you don't want to build UI ## Build your own admin panel Headless mode: the full engine behind your own UI. Everything it computes is reachable from a TypeScript SDK and REST API — query it, or drop the whole module in. ```typescript const report = await ss.reports.attribution({ workspace: "acme-retail", model: "predictive", date_range: { from: "2026-06-01", to: "2026-06-30", }, group_by: ["channel"], }); ``` Response (200 OK · 2.4 GB scanned · 2.4 DQU): ```json { "rows": [ { "channel": "Paid Search", "spend": 48210, "roas": 4.02 }, { "channel": "Paid Social", "spend": 31450, "roas": 2.77 } ] } ``` Embedding the white-label module: ```typescript // server: mint a scoped embed token const token = await ss.embedTokens.create({ workspace: "acme-retail", }); // client: render the white-label module ``` API preview — final shapes may change before general availability. Every workspace ships with MCP access, so AI agents can query the engine directly. Prefer config-as-code in your own warehouse? See the [CLI](https://segmentstream.com/cli.md). ## Pay for what you process A platform fee per plan, plus two meters. No per-seat pricing, no per-client surprises. - **DPU — Data Processing Unit.** Event ingestion, identity resolution and model refreshes — every unit of data the engine processes on your behalf. - **DQU — Data Query Unit.** Reports, API calls and AI-agent queries, metered by data scanned. One DQU is one gigabyte. Each plan includes monthly DPU and DQU quotas — unit sizing and early-access pricing are shared with the waitlist first. ## Join the waitlist SegmentStream Cloud is rolling out to a small group of design partners. Join the waitlist at https://segmentstream.com/cloud to get early pricing. --- # Marketing Measurement Engine Twelve marketing measurement whitepapers covering attribution, incrementality, predictive modeling, and budget allocation. Every model explained — no black boxes. Most vendors describe their product with marketing slogans. SegmentStream describes the actual algorithms. Read the full technical pages below, or use this index to understand the scope of the engine. ## Capabilities - [Identity Graph](https://segmentstream.com/measurement-engine/identity-graph.md): Deterministic identity stitching across devices and touchpoints — login events, email matching, CRM joins — unified into a single user profile in BigQuery. The foundational data layer that cross-channel attribution and scoring depend on. - [Cross-Channel Attribution](https://segmentstream.com/measurement-engine/cross-channel-attribution.md): First-click attribution and predictive conversion maturation — two methods that cross-validate to produce click-time revenue you can optimize against before all conversions are reported. - [Predictive Cross-Channel Attribution](https://segmentstream.com/measurement-engine/predictive-attribution.md): ML-powered projection that fills the click-time attribution gap. See projected ROAS within days, not weeks — while cohorts are still maturing. - [Self-Reported Reattribution](https://segmentstream.com/measurement-engine/self-reported-reattribution.md): LLM classification of free-text survey responses, stitched to sessions via BigQuery. Measures channels traditional tracking can't see — word of mouth, podcasts, offline. - [CRM Funnel Attribution](https://segmentstream.com/measurement-engine/crm-funnel-attribution.md): Connect CRM, ERP, or any BigQuery table as a conversion source. Track every funnel stage from MQL to Closed/Won — the ground truth that calibrates predictive models and reconciles true campaign ROI. - [Marginal Analytics](https://segmentstream.com/measurement-engine/marginal-analytics.md): Diminishing returns curve fitting that models marginal ROAS per channel at every spend level. See exactly where each dollar stops working, calculate the optimal budget mix, and get specific reallocation recommendations. - [Automated Budget Allocation](https://segmentstream.com/measurement-engine/automated-budget-allocation.md): One-click execution of budget recommendations across every ad platform. Applies changes from Marginal Analytics, tracks prediction accuracy, and refines the model with reinforced learning. - [Incrementality Testing](https://segmentstream.com/measurement-engine/incrementality.md): Geo holdout experiments with synthetic control groups. Three-phase market selection, A/A validation, and statistical rigor that measures what your ads actually add. - [Server-Side Conversion Tracking](https://segmentstream.com/measurement-engine/server-side-tracking.md): Stop in-page pixels from broadcasting your audience to competitors. Forward only qualifying conversions to each ad platform server-to-server, with the hashed CRM identifiers that unlock cross-device advanced matching. --- # SegmentStream CLI An open-source CLI for warehouse-native marketing attribution. Connect BigQuery, Snowflake, or Databricks, add any data source, and run the full Measurement Engine locally — your data never leaves your warehouse. The SegmentStream CLI is in private beta. Install: ```bash curl -fsSL https://segmentstream.com/cli/install.sh | sh ``` ## How it works ### 01 — Initialize Run `segmentstream init` in any repository. It scaffolds a `config.yaml` and the folders where your connectors and transformations live — version-controlled, reviewable, and entirely yours. ### 02 — Connect a warehouse BigQuery, Snowflake, or Databricks. SegmentStream reads and writes inside your warehouse — it never copies or moves your data out of it. ### 03 — Add any data source Ad platforms, website and app events, your CRM — start with built-in connectors, or add your own. Describe the source you need and it gets built for you. ### 04 — Configure attribution Choose your attribution model, set the identity-graph keys that stitch users into one journey, and define what counts as a conversion. It all lives in one config. ### 05 — Run One command runs the full pipeline — identity stitching, attribution, and reporting — and writes the results straight back to your warehouse. ## The full Measurement Engine, on your terminal Every model the platform runs is available from the CLI — Identity Graph, Cross-Channel Attribution, Predictive Attribution, Self-Reported Reattribution, CRM Funnel Attribution, Marginal Analytics, Automated Budget Allocation, Incrementality Testing, and Server-Side Tracking — explained in full, with no black boxes. --- # Integrations Connects to your entire marketing and AI stack. Full-funnel measurement and optimization across every channel. From ad spend to real revenue and LTV. Accurate, unified, actionable. ## Analytics Platforms You already track events and conversions across web and app — using tools like GA4, Adobe Analytics, Heap, or Amplitude. SegmentStream layers advanced measurement on top: identity graph, cross-channel attribution, budget optimization and more. Less code. More insights. ## Built on top of the analytics you already run. - [Google Analytics 4](https://segmentstream.com/integrations/google-analytics-4.md): Ingest events, conversions, and audiences from your primary analytics. - [Adobe Analytics](https://segmentstream.com/integrations/adobe-analytics.md): Connect enterprise Adobe event streams to attribution. - [Heap by Contentsquare](https://segmentstream.com/integrations/heap-contentsquare.md): Import auto-captured events and funnels for measurement. - [Amplitude](https://segmentstream.com/integrations/amplitude.md): Bring product-analytics events into cross-channel measurement. - [Mixpanel](https://segmentstream.com/integrations/mixpanel.md): Enrich Mixpanel event data with paid-media attribution. - [PostHog](https://segmentstream.com/integrations/posthog.md): Open-source product analytics, attribution-ready. - [Twilio Segment](https://segmentstream.com/integrations/twilio-segment.md): Customer data platform for event pipelines. - [Snowplow](https://segmentstream.com/integrations/snowplow.md): Open-source behavioral data pipeline. - [RudderStack](https://segmentstream.com/integrations/rudderstack.md): Warehouse-first customer data platform. ## Mobile install and in-app conversion data. Ingest mobile measurement from iOS and Android apps. Install attribution, in-app events, and post-install revenue flow into SegmentStream alongside web conversions — unified across devices. - AppsFlyer: Mobile measurement for iOS and Android apps. - Adjust: Mobile attribution and in-app analytics. ## Every paid media channel, in one place. All your cost, click, and impression data — pulled from every ad platform, deduplicated and unified. Analyze and optimize cross-channel performance from a single view, without switching between dozens of platforms and ad accounts. - Google Ads: Search, Display, YouTube, Performance Max. - Meta Ads: Facebook, Instagram, Audience Network. - TikTok Ads: Performance and brand campaigns. - LinkedIn Ads: B2B paid social and lead gen. - Microsoft Ads: Bing Search and Microsoft Audience Network. - Snapchat Ads: Direct-response and awareness campaigns. - X Ads: Paid posts and brand activations. - Pinterest Ads: Shopping and inspiration ads. - Reddit Ads: Community-driven paid reach. - Display & Video 360: Programmatic display, video, and CTV. - Campaign Manager 360: Ad serving and campaign reporting. - AdRoll: Retargeting and prospecting. - Criteo: Commerce media and retargeting. - RTB House: Deep-learning retargeting. - Xandr: Programmatic DSP (Microsoft). - Awin: Global affiliate and partnerships network. - Impact: Partnerships and affiliate tracking. - Rakuten: Global affiliate and partnerships network. - The Trade Desk: Programmatic DSP for display, video, CTV, audio. - StackAdapt: Multichannel programmatic advertising. - Adform: European DSP and ad-serving platform. - Taboola: Native content discovery ads. - Outbrain: Native content recommendation ads. - Applovin: Mobile gaming and app install ads. ## Your AI tools. Our measurement brain. Give AI agents direct access to cross-channel attribution, measurement insights, and budget recommendations — every SegmentStream capability is exposed as a tool your AI agent can call. Ask any questions, get answers in seconds, instantly create custom reports. Directly from your AI. - Claude: Anthropic's conversational AI. - Claude Cowork: Anthropic's collaborative AI workspace. - Claude Code: Anthropic's agentic coding CLI. - Cursor: AI-first code editor. - ChatGPT: OpenAI's conversational AI. - Codex: OpenAI's agentic coding CLI. - Gemini: Google's conversational AI. - Perplexity: AI-powered answer engine. - Microsoft Copilot: Microsoft's AI assistant for work. - Replit: AI-powered cloud coding environment. - Lovable: AI app and web builder. - Bolt.new: AI full-stack web app builder. - Windsurf: AI-first code editor. - Any MCP client: Bring your own MCP-compatible AI tool. ## Works with your marketing operations stack. Bid management, creative automation, fraud protection, feed optimization, ad verification — whatever you run to operate and optimize paid media, SegmentStream plugs into the same workflow. These tools shape the inputs. SegmentStream measures the outcomes. - Opteo: Google Ads optimization for agencies. - Lunio: PPC click fraud protection. - Smartly: Paid social creative automation. - Skai: Search, social, and retail ad management. - Celtra: Creative automation at scale. - Productsup: Product feed management for commerce ads. - DoubleVerify: Ad verification and viewability measurement. - Integral Ad Science: Ad verification and brand safety. - CHEQ: Go-to-market security and ad fraud protection. - Hightouch: Warehouse-native data activation. - Wunderkind: Onsite visitor identification and conversion. - LiveRamp: Enterprise identity resolution. ## Measurement and optimization toward high-value customers. Orders, subscriptions, and billing data pulled from your commerce and payment platforms. Measure marketing on real revenue and lifetime value — not just first-purchase orders. - Shopify: Orders, customers, and subscription data. - Stripe: Subscription revenue, billing, and LTV. ## Every ad dollar tied to real pipeline and closed revenue. Opportunities, deals, and accounts pulled from your CRM and marketing automation — reconciled with the paid touchpoints that drove them. Measure marketing on revenue, not MQLs. - Salesforce: Opportunities, accounts, and enterprise revenue. - HubSpot: Deals, pipeline, and marketing contacts. - Marketo: B2B marketing automation (Adobe). - Microsoft Dynamics 365: Enterprise CRM and business apps. ## Customer journey tracking, enhanced by your CDP data. Customer profiles and identifiers from your CDP feed directly into SegmentStream's identity graph, linking every touchpoint across devices and sessions into one continuous journey. Better stitching means more accurate attribution. - Braze: Customer engagement and messaging events. - Klaviyo: Email and SMS revenue for commerce. - Bloomreach: Commerce experience and engagement platform. ## Every phone conversation, tracked as a conversion. Capture inbound-call conversions from every channel and tie them back to the ad, campaign, and keyword that drove them. Call data flows into the same attribution model as web and app conversions. - CallRail: Call tracking for SMB and mid-market. - Invoca: Enterprise conversation intelligence. - Infinity: Enterprise call tracking and analytics. ## SegmentStream data, inside your existing BI stack. All SegmentStream data lives in your data warehouse — cross-channel ROAS, pipeline, LTV, attribution, cost. Connect any BI tool your team already uses and get analytics-ready data without custom SQL or data prep work. - Looker: Google's semantic-layer BI. - Power BI: Microsoft's enterprise BI platform. - Tableau: Salesforce's visual analytics platform. - Qlik: Enterprise BI with associative engine. - Domo: Cloud BI for executives. - ThoughtSpot: Search-based analytics. - Sigma: Spreadsheet-style cloud BI. - Hex: Collaborative data notebooks. - Omni: Modern BI with embedded semantic layer. ## Any business data, from your warehouse. Connect any data warehouse or database to pull the business context that lives outside your CRM or ad platforms. Profit margins, stock levels, product attributes, customer segments — all become available for attribution, measurement, and paid media optimization. - Google BigQuery: Google Cloud data warehouse. - Snowflake: Cloud data platform. - Amazon Redshift: AWS data warehouse. - Databricks: Lakehouse platform. - Microsoft Azure: Azure cloud data services. - ClickHouse: Open-source columnar warehouse. - PostgreSQL: Open-source relational database. - Supabase: Open-source Postgres data platform. - SAP Data Warehouse: Enterprise data warehouse cloud. ## Don't see an integration you need? Not a problem. If your platform has an API, we can build a new connector quickly — connectors ship fast based on demand. And if there's no API, there's always a path: ad platforms can export their data into a Google Sheet and we import it from there; CRMs can export to a data warehouse and we pull it from there, or send data over webhooks. --- # Pricing SegmentStream is a high-touch measurement partner: an AI-powered platform plus experts in marketing science, incrementality, and attribution. Every engagement includes dedicated onboarding, ongoing consulting, and enterprise-level support, scoped and priced after discovery. Scoped to your business, priced after discovery. We learn about your business and goals first, then propose the capabilities and engagement structure that fit. ## Everything in the platform Which capabilities you use is decided together during scoping. ### Measurement engine - [Identity Graph](https://segmentstream.com/measurement-engine/identity-graph.md) - [Cross-Channel Attribution](https://segmentstream.com/measurement-engine/cross-channel-attribution.md) - [Predictive Cross-Channel Attribution](https://segmentstream.com/measurement-engine/predictive-attribution.md) - [Self-Reported Reattribution](https://segmentstream.com/measurement-engine/self-reported-reattribution.md) - [CRM Funnel Attribution](https://segmentstream.com/measurement-engine/crm-funnel-attribution.md) - [Marginal Analytics](https://segmentstream.com/measurement-engine/marginal-analytics.md) - [Automated Budget Allocation](https://segmentstream.com/measurement-engine/automated-budget-allocation.md) - [Incrementality Testing](https://segmentstream.com/measurement-engine/incrementality.md) - [Server-Side Conversion Tracking](https://segmentstream.com/measurement-engine/server-side-tracking.md) ### AI & platform - MCP access for Claude, Cursor, and any MCP client - AI measurement skills - Historical backfill - Marketing Measurement course access ### Security & compliance - Advanced security assessment - DPA and GDPR compliance - Data residency (US/EU) - Custom retention policy - Uptime SLA ### Support - Dedicated measurement expert - Priority SLA and response times - Founder-led consulting - Custom MSA and DPA ## Every engagement includes Dedicated onboarding and on-going expert guidance. ### Tailored setup Not just a platform login and link to documentation. Our team will take full care of your setup and team onboarding, customized to your exact stack and challenges. You get a solution, not a tool. ### Ongoing consulting Dedicated communication with our senior measurement experts, helping you gain maximum value from SegmentStream measurement and optimization engine. ### Enterprise-level support Priority response with agreed SLAs, security review, DPA and GDPR compliance, US or EU data residency, and custom terms when your procurement needs them. ## How pricing works One discovery call, then one proposal. 1. **Discovery.** A working session on how you sell: where conversions happen (online, in a CRM, offline), which platforms you run, what your data stack looks like, and which questions your team needs answered. 2. **Scoping.** We propose the capabilities that actually fit and price the engagement to that scope: channels, data volume, and how much expert time you need. One proposal covers platform, onboarding, consulting, and support. 3. **Engagement.** Onboarding starts with your dedicated expert. Pricing is per project, with no seat cap, so your whole team works from the same numbers. ## Frequently asked questions ### Why don't you publish plans and prices? Because the right scope varies a lot between clients. An e-commerce brand running five ad platforms and a B2B company closing deals in a CRM need different capabilities, different onboarding, and different amounts of expert time. Pricing that after a discovery call is more honest than a plan tier that fits neither. ### Is the pricing per user? No. Pricing is per project, with no seat cap and no per-user charge. Price scales with scope: channels, data volume, and the level of expert involvement. ### What happens in the discovery call? We map how you sell and where conversions happen, look at your ad platforms and data stack, and agree on the questions measurement should answer. After the call you get a written proposal with the capabilities we recommend and the price for that scope. ### What does onboarding look like? Your dedicated expert does the setup with you: data sources, identity graph, conversions, attribution models, and the first dashboards. You review the numbers together before the team starts using them. ### Can we negotiate the MSA? Our standard [MSA](https://segmentstream.com/legal/msa.md) and [DPA](https://segmentstream.com/dpa.md) are published. Redlines, a custom DPA, and security questionnaire responses are handled during scoping. We do not sign BAAs. ### What AI tools work with SegmentStream? SegmentStream connects via Model Context Protocol (MCP). It works with Claude Code, Claude Cowork, Cursor, and any MCP-compatible AI client. You use your own AI subscription — SegmentStream provides the measurement infrastructure your AI connects to. ## Let's talk Book a discovery call with our senior measurement partner. Get in touch via the chat on [segmentstream.com/pricing](https://segmentstream.com/pricing). ## Related Pages - [Measurement Engine](https://segmentstream.com/measurement-engine.md) - [Integrations](https://segmentstream.com/integrations.md) - [For Agencies](https://segmentstream.com/agency.md) --- # Marketing Measurement for Agencies & Consultants AI-native marketing analytics, built for agencies of a new era. Answer any client's data question in seconds. Generate branded reports on demand. Scale measurement across every account through AI. ## One AI. All Your Clients. Ask about any client. Get cross-channel answers in seconds. Generate white-label reports on demand. No analyst queue, no manual pulls — just AI and SegmentStream. ## Client Deliverables Your AI generates analyst-grade reports powered by SegmentStream's measurement engine. Brand them, share them, and bill for them. - Campaign performance reports - Root-cause investigations - Budget reallocation plans - Month-over-month analysis ## Agency Economics - Clients per account manager can grow from 5-8 to 30-50. - Data questions can be answered in seconds instead of analyst queues. - Monthly reporting prep becomes automated. - Budget optimization can run continuously instead of weekly at best. ## Measurement Capabilities Attribution, optimization, and incrementality. Trusted by leading enterprises and agencies. $100M+ in managed ad spend. - [Measurement Engine](https://segmentstream.com/measurement-engine.md) - [Pricing](https://segmentstream.com/pricing.md) --- # Blog Articles, product updates, and insights on marketing measurement, attribution, and budget optimization. ## Categories - [Articles](https://segmentstream.com/blog/articles.md) - [Product Updates](https://segmentstream.com/blog/product-updates.md) - [Company News](https://segmentstream.com/blog/company-news.md) ## Posts - [Introducing the SegmentStream MCP Server](https://segmentstream.com/blog/product-updates/introducing-segmentstream-mcp-server.md): A way for AI assistants to directly connect to SegmentStream's measurement and optimization engine — and take action through it. - [Incrementality Measurement Guide (2026)](https://segmentstream.com/blog/articles/incrementality-measurement-guide.md): This guide explores the methodologies behind incrementality measurement and dives into the pros and cons of the approach. - [Marketing Attribution 101 (2026 Guide)](https://segmentstream.com/blog/articles/marketing-attribution-101.md): In this article, we're digging deeper into the topic of marketing attribution: benefits, attribution models, and typical challenges. - [Identity Graph: The Foundation of Attribution](https://segmentstream.com/blog/articles/what-is-identity-graph-the-foundation-of-marketing-attribution.md): Identity graphs are the foundation for accurate marketing attribution. Discover how they connect user journeys across devices and browsers. - [Click Propagation and Attribution Accuracy](https://segmentstream.com/blog/articles/what-is-click-propagation-and-how-it-impacts-marketing-attribution-accuracy.md): Browser switches and shared links break marketing attribution. Learn how click propagation technology restores accuracy across devices and platforms. - [What Is Conversion Maturation in Attribution?](https://segmentstream.com/blog/articles/what-is-conversion-maturation-in-marketing-attribution.md): Conversion Maturation is the hidden delay between ad clicks and reported conversions. Learn how it impacts attribution accuracy and how to fix it. - [The Truth About Next-Gen MMM](https://segmentstream.com/blog/articles/truth-next-gen-mmm-what-marketers-need-know.md): Next-gen MMM tools promise better results than traditional media mix models. Here's what marketers really need to know before investing. - [The Misuse of Geo-Holdout Tests](https://segmentstream.com/blog/articles/misuse-geo-holdout-tests-guide-non-technical-leaders.md): “Geo-holdout tests promise to reveal true incremental ROAS, but they often oversell precision. A practical guide for non-technical marketing leaders.” - [LTV-Based Ads Optimization Done Right](https://segmentstream.com/blog/articles/implementing-ltv-based-ads-optimization-right-way.md): LTV-based ads optimisation is challenging but essential for DTC, SaaS, and subscription businesses. Learn how to implement it the right way. - [Synthetic Conversions for Upper-Funnel Wins](https://segmentstream.com/blog/articles/synthetic-conversions-your-secret-weapon-upper-funnel-wins.md): “Learn how synthetic conversions help ad platforms optimise upper-funnel campaigns by turning engaged website visits into conversion signals.” - [Debunking MMM Myths: A Waste of Money?](https://segmentstream.com/blog/articles/debunking-mmm-myths-why-its-waste-money.md): Marketing Mix Modeling doesn't work for 99% of online businesses. Learn why MMM fails most companies and what measurement approach actually works. - [Why to Avoid Target ROAS Bidding](https://segmentstream.com/blog/articles/why-you-should-avoid-using-target-roas-bidding.md): Relying on Average ROAS and CPA can burn your ad budget. Learn why Target ROAS and Target CPA bidding strategies often do more harm than good. - [Why Geo-Lift Testing Falls Short](https://segmentstream.com/blog/articles/why-geo-lift-testing-falls-short-measure-true-ads-incrementality.md): Geo-lift testing promises unbiased incrementality measurement, but it hides subtle pitfalls. Learn why it falls short and what works better. - [Brand Awareness = Low-Quality Targeting?](https://segmentstream.com/blog/articles/brand-awareness-fancy-term-targeting-low-quality-audience.md): Brand awareness campaigns sound impressive but often target low-quality audiences. Learn why this common strategy wastes your ad budget. - [How to Validate Your Attribution Model?](https://segmentstream.com/blog/articles/how-validate-your-attribution-model.md): A step-by-step methodology to find and validate the best attribution model for your business and maximise marketing mix revenue. - [Debunking Post-View Attribution](https://segmentstream.com/blog/articles/debunking-post-view-attribution.md): Post-view attribution claims to measure the silent influence of display ads. Here's why it's misleading and what smart marketers should use instead. - [SegmentStream rated #1 B2C Attribution by G2](https://segmentstream.com/blog/company-news/number-one-in-b2-attribution-by-g2.md): SegmentStream rated #1 in B2C Attribution and Custom Attribution on G2 Summer Reports. See why teams choose SegmentStream for accurate, AI-powered measurement. - [Marketing Attribution Challenges & Solutions](https://segmentstream.com/blog/articles/marketing-attribution-common-challenges.md): Learn about the top 3 challenges modern marketers face in marketing attribution due to tracking restrictions and how to overcome them.. - [SegmentStream is officially a Meta Business Partner!](https://segmentstream.com/blog/company-news/segmentstream-meta-business-partner.md): “SegmentStream is now an official Meta Business Partner. The Badged status recognises our product innovation and service excellence.” - [SegmentStream: Official Google Cloud Partner](https://segmentstream.com/blog/company-news/gcp-partnership.md): SegmentStream is now a certified Google Cloud Platform partner. Learn how this GCP partnership strengthens our AI-powered marketing measurement platform. --- # Blog: Articles Articles from SegmentStream. ## Posts - [Incrementality Measurement Guide (2026)](https://segmentstream.com/blog/articles/incrementality-measurement-guide.md): This guide explores the methodologies behind incrementality measurement and dives into the pros and cons of the approach. - [Marketing Attribution 101 (2026 Guide)](https://segmentstream.com/blog/articles/marketing-attribution-101.md): In this article, we're digging deeper into the topic of marketing attribution: benefits, attribution models, and typical challenges. - [Identity Graph: The Foundation of Attribution](https://segmentstream.com/blog/articles/what-is-identity-graph-the-foundation-of-marketing-attribution.md): Identity graphs are the foundation for accurate marketing attribution. Discover how they connect user journeys across devices and browsers. - [Click Propagation and Attribution Accuracy](https://segmentstream.com/blog/articles/what-is-click-propagation-and-how-it-impacts-marketing-attribution-accuracy.md): Browser switches and shared links break marketing attribution. Learn how click propagation technology restores accuracy across devices and platforms. - [What Is Conversion Maturation in Attribution?](https://segmentstream.com/blog/articles/what-is-conversion-maturation-in-marketing-attribution.md): Conversion Maturation is the hidden delay between ad clicks and reported conversions. Learn how it impacts attribution accuracy and how to fix it. - [The Truth About Next-Gen MMM](https://segmentstream.com/blog/articles/truth-next-gen-mmm-what-marketers-need-know.md): Next-gen MMM tools promise better results than traditional media mix models. Here's what marketers really need to know before investing. - [The Misuse of Geo-Holdout Tests](https://segmentstream.com/blog/articles/misuse-geo-holdout-tests-guide-non-technical-leaders.md): “Geo-holdout tests promise to reveal true incremental ROAS, but they often oversell precision. A practical guide for non-technical marketing leaders.” - [LTV-Based Ads Optimization Done Right](https://segmentstream.com/blog/articles/implementing-ltv-based-ads-optimization-right-way.md): LTV-based ads optimisation is challenging but essential for DTC, SaaS, and subscription businesses. Learn how to implement it the right way. - [Synthetic Conversions for Upper-Funnel Wins](https://segmentstream.com/blog/articles/synthetic-conversions-your-secret-weapon-upper-funnel-wins.md): “Learn how synthetic conversions help ad platforms optimise upper-funnel campaigns by turning engaged website visits into conversion signals.” - [Debunking MMM Myths: A Waste of Money?](https://segmentstream.com/blog/articles/debunking-mmm-myths-why-its-waste-money.md): Marketing Mix Modeling doesn't work for 99% of online businesses. Learn why MMM fails most companies and what measurement approach actually works. - [Why to Avoid Target ROAS Bidding](https://segmentstream.com/blog/articles/why-you-should-avoid-using-target-roas-bidding.md): Relying on Average ROAS and CPA can burn your ad budget. Learn why Target ROAS and Target CPA bidding strategies often do more harm than good. - [Why Geo-Lift Testing Falls Short](https://segmentstream.com/blog/articles/why-geo-lift-testing-falls-short-measure-true-ads-incrementality.md): Geo-lift testing promises unbiased incrementality measurement, but it hides subtle pitfalls. Learn why it falls short and what works better. - [Brand Awareness = Low-Quality Targeting?](https://segmentstream.com/blog/articles/brand-awareness-fancy-term-targeting-low-quality-audience.md): Brand awareness campaigns sound impressive but often target low-quality audiences. Learn why this common strategy wastes your ad budget. - [How to Validate Your Attribution Model?](https://segmentstream.com/blog/articles/how-validate-your-attribution-model.md): A step-by-step methodology to find and validate the best attribution model for your business and maximise marketing mix revenue. - [Debunking Post-View Attribution](https://segmentstream.com/blog/articles/debunking-post-view-attribution.md): Post-view attribution claims to measure the silent influence of display ads. Here's why it's misleading and what smart marketers should use instead. - [Marketing Attribution Challenges & Solutions](https://segmentstream.com/blog/articles/marketing-attribution-common-challenges.md): Learn about the top 3 challenges modern marketers face in marketing attribution due to tracking restrictions and how to overcome them.. --- # Blog: Product Updates Product Updates from SegmentStream. ## Posts - [Introducing the SegmentStream MCP Server](https://segmentstream.com/blog/product-updates/introducing-segmentstream-mcp-server.md): A way for AI assistants to directly connect to SegmentStream's measurement and optimization engine — and take action through it. --- # Blog: Company News Company News from SegmentStream. ## Posts - [SegmentStream rated #1 B2C Attribution by G2](https://segmentstream.com/blog/company-news/number-one-in-b2-attribution-by-g2.md): SegmentStream rated #1 in B2C Attribution and Custom Attribution on G2 Summer Reports. See why teams choose SegmentStream for accurate, AI-powered measurement. - [SegmentStream is officially a Meta Business Partner!](https://segmentstream.com/blog/company-news/segmentstream-meta-business-partner.md): “SegmentStream is now an official Meta Business Partner. The Badged status recognises our product innovation and service excellence.” - [SegmentStream: Official Google Cloud Partner](https://segmentstream.com/blog/company-news/gcp-partnership.md): SegmentStream is now a certified Google Cloud Platform partner. Learn how this GCP partnership strengthens our AI-powered marketing measurement platform. --- # Glossary Plain-English definitions for marketing measurement: attribution models, incrementality testing, marketing mix modeling, identity graphs, and the metrics that actually drive budget decisions. ## Entries - [Incrementality Testing vs Attribution](https://segmentstream.com/glossary/incrementality-testing-vs-attribution-modelling.md): Comparing Incrementality testing vs Attribution modelling to explore pros and cons of each solution for analytics and optimization. - [Incrementality Testing vs Media Mix Modeling](https://segmentstream.com/glossary/incrementality-testing-vs-media-mix-modelling.md): Exploring Incrementality testing vs. Media Mix Modeling (MMM): Understand the benefits and challenges of each for media mix optimization. - [Marketing Mix Modeling vs Attribution](https://segmentstream.com/glossary/marketing-mix-modeling-vs-attribution.md): What's the difference between MMM and Attribution modelling? What are the pros and cons of both approaches and is there an all-in-one alternative? - [What Is Attribution Modeling?](https://segmentstream.com/glossary/attribution-modelling.md): Attribution modeling assigns credit to marketing touchpoints that drive conversions. Learn model types, how to choose, and where attribution breaks down. - [What Is Incrementality Testing?](https://segmentstream.com/glossary/incrementality-testing.md): Explore how incrementality testing works, discover three methods to measure it, and understand its difference from A/B tests. - [What Is Marketing Attribution?](https://segmentstream.com/glossary/marketing-attribution.md): Marketing attribution identifies which channels drive conversions. Learn attribution models, how they work, and how to choose the right one. - [What Is Marketing Mix Optimization?](https://segmentstream.com/glossary/marketing-mix-optimization-tools-and-challenges.md): Marketing mix optimization challenges — ad platform bias, attribution limits, incrementality testing gaps, and MMM shortcomings explained. - [What Is Media Mix Modeling (MMM)?](https://segmentstream.com/glossary/media-mix-modelling.md): What is Media Mix Modelling (MMM)? Prerequisites, strengths, limitations, alternatives, and how it compares to attribution. - [What Is Multi-Touch Attribution?](https://segmentstream.com/glossary/multi-touch-attribution.md): Multi-touch attribution distributes conversion credit across every touchpoint in the customer journey. Learn how MTA models work and when to use each one. - [What Is Predictive Attribution?](https://segmentstream.com/glossary/predictive-attribution.md): What's the difference between predictive and retrospective attribution and how does it work? read this article to find out! --- # Podcasts Podcast episodes and conversations about marketing measurement, attribution, incrementality, and AI-powered optimization. ## Episodes - [DTC Podcast: The Hard Truth About Marketing Measurement](https://segmentstream.com/resources/podcasts/podcast-the-hard-truth-about-marketing-measurement.md): Listen to the latest episode of the DTC Podcast, featuring Constantine Yurevich, founder of SegmentStream, as he delivers the most comprehensive breakdown of the marketing measurement landscape ever recorded. - [What DTC Brands Get Wrong About Attribution](https://segmentstream.com/resources/podcasts/what-dtc-brands-get-wrong-about-attribution.md): In this episode of Chew on This, Ron and Ash are joined by Constantine Yurevich, Founder of SegmentStream, to discuss the challenges and future of marketing measurement and attribution. If you care about scaling profitably, avoiding wasted ad spend, and finally gaining clarity on what’s actually driving results — this one’s for you. - [Attribution for SaaS & Lead-Gen Businesses](https://segmentstream.com/resources/podcasts/how-to-solve-marketing-attribution-for-saas-and-lead-gen-businesses.md): In this interview, Oren Greenberg and Constantine Yurevich discuss attribution problems in SaaS and PLG marketing. They explain how SegmentStream uses predictive lead scoring and LTV modeling for cookie-less attribution and optimizing ad platforms with real-time value signals to boost marketing ROI. - [DTC Podcast: Marginal ROAS & How Meta’s Algorithm Really Works](https://segmentstream.com/resources/podcasts/dtc-podcast-marginal-roas.md): In this episode of the DTC Podcast, host Eric Dyck welcomes back Constantine Yurevich, founder of SegmentStream, for a tactical conversation on modern marketing measurement and how to truly optimize media budgets — especially on Meta. - [Why Attribution Is Broken — And What Comes Next](https://segmentstream.com/resources/podcasts/jordan-west-podcast-why-attribution-is-broken.md): In this episode, Jordan West talks with Constantine Yurevich, Founder & CEO of SegmentStream, about the shortcomings of cookie-based tracking and traditional attribution methods in today’s privacy-first landscape. They explore why multi-touch attribution and geo holdouts often fail, what defines a modern marketing intelligence platform, and how brands can make better budget decisions despite having incomplete data. - [DTC Podcast: Why Attribution Is Dead — And What To Do Instead](https://segmentstream.com/resources/podcasts/dtc.md): In the recent DTC podcast, SegmentStream’s CEO discusses multi-device customer experiences and why marketing attribution is dead. --- # About SegmentStream The company, the principles, and how we work. ## The Work Marketing measurement is closer to critical thinking than to marketing. It is the architecture of tracking, the discipline of data consolidation, and the math that turns ad spend into decisions a finance team can defend. It is also the processes and decision frameworks embedded into a marketing team that turn those decisions into moved budget. We started SegmentStream in 2018. The category had drifted into a belief system, where teams trusted attribution models and MMM frameworks the way people trust horoscopes. Confident. Detailed. Mostly fiction. The brands paying for the ads deserved better. Eight years on, the work has grown into the [SegmentStream Measurement Engine](https://segmentstream.com/measurement-engine): cross-channel attribution, identity stitching, incrementality experiments, marginal analytics, automated budget allocation, and AI agents that turn raw marketing data into answers a finance team can defend. We built it for the brands paying for the ads. Independent of the platforms we measure, grounded in evidence, published in full, wired to act, delivered as a system, taught from first principles. That is the only side we work for. ## Our Principles ### 01. Independent of the platforms we measure. *We work for the advertiser. That is the only side we work for.* We do not co-market with ad platforms. No joint case studies. No subsidised client placements. No "strategic partnerships." No referral kickbacks. When you read a SegmentStream report, no ad platform has been in the room shaping the conclusion. The same way an auditor's only client is the company whose books they audit, our only client is the advertiser. Our incentive is to count honestly, because the advertiser is the one paying us to count. ### 02. Evidence over modeling. *Deterministic before modeled. Always labeled.* A deterministic measurement is a fact. A model is a hypothesis. Until evidence backs it, we label it that way in the report. Our default is deterministic data the customer owns. We label every modeled number as modeled. We run incrementality experiments when the answer matters. We report the confidence interval, not just the headline lift. When the data does not support a conclusion, we say so. Platform-reported conversions get the same treatment. We collect them, but we never let them lead. ### 03. No black boxes. *Every model published. Every number auditable.* Every model we ship is one the customer's finance team can audit. We publish how attribution credit is assigned. We publish how identity stitching works. We publish how marginal analytics evaluates each next dollar of spend. We publish the geo-holdout methodology, the priors, the confidence intervals, and the parts where the method has known limits. The [SegmentStream Measurement Engine](https://segmentstream.com/measurement-engine) ships as nine open whitepapers, each one a full explanation of one method. If you cannot find the math behind a number we report, that is a bug. Flag it and we will publish it. A measurement number you cannot defend is a measurement number you cannot use. ### 04. Built for action. *Measurement is only useful when it changes a decision.* A measurement report that ends in a slide deck and stays with the analyst does not change a business. Numbers get pulled, decks get written, the marketing lead nods, and the budget moves the way it was already going to move. SegmentStream is wired to act on what it measures. Marginal analytics says where the next dollar should go. [Automated budget allocation](https://segmentstream.com/measurement-engine/automated-budget-allocation) moves the budget against that signal, daily. Incrementality experiments validate that the move worked. The measurement, the recommendation, and the execution sit in one system. Reports still exist, but they read as exhaust — what falls out of decisions made and validated, not what the team has to read to figure out what to do. ### 05. A system, not a tool. *Technology, process, expertise. All three, together.* Marketing measurement is technology, process, and expertise working as one system. The technology runs the math and the data infrastructure that feeds it. The process is how the company organizes its decisions around what the system surfaces. The expertise is the methodology knowledge that connects them. Skip any one and the system breaks at that layer. The "just add AI" version of this skips the harder work. A chatbot on top of a dashboard does not change how a team answers questions or moves a budget. The operating system has to change — the process, the methodology, the people running it. SegmentStream ships as a full system. The measurement engine handles the math. We help wire the measurement into how decisions actually get made each week. Our experts work alongside the team running the budget. The value comes from all three working together. ### 06. Education over selling. *Teach the discipline before selling the tool.* Marketing measurement is a discipline. It becomes useful when the buyer understands the math. So we teach the discipline in long form. The [nine measurement engine whitepapers](https://segmentstream.com/measurement-engine) explain each method from first principles. The [9.5-hour course on modern marketing measurement](https://segmentstream.com/course) teaches the discipline end-to-end. Both are open. Both are the work a buyer needs to do before the platform becomes useful. When a client buys SegmentStream after working through the methodology, they buy for the right reasons and they stay. So we write the textbook first. ## Where to go from here - [The measurement engine](https://segmentstream.com/measurement-engine) — Nine whitepapers, every model explained. - [The course](https://segmentstream.com/course) — 9.5 hours, taught from first principles. - [Trust & security](https://segmentstream.com/trust) — How we handle your customer data and who has access. - [Get in touch](https://segmentstream.com/pricing) — Request access to the SegmentStream platform. --- # M&A Interest > Information for strategic acquirers and corporate development teams. SegmentStream is selectively open to acquisition conversations where the combined offering beats the sum of its parts. This page is for strategic acquirers and corporate development teams. SegmentStream is selectively open to acquisition proposals from strategic buyers positioned to scale the technology, team, and category position we've built. We're prioritising conversations where the combined offering is stronger than the sum of its parts. To get in touch, reach out to the founders directly on LinkedIn or X (linked below). Acquisition conversations go straight to the founders, not into a sales pipeline. Trusted by leading brands across verticals — E-Commerce, B2B SaaS, Finance, Travel, Energy, Automotive, and more. ## About SegmentStream An AI-native marketing intelligence platform that helps brands measure and optimise their cross-channel marketing investment. SegmentStream is an independent marketing measurement company, operating across marketing analytics, cross-channel attribution, marketing mix optimization, incrementality testing, paid media intelligence, and agentic AI ad optimization. Explore the [Measurement Engine](https://segmentstream.com/measurement-engine.md) for every capability, explained. Brands turn to SegmentStream when out-of-the-box analytics (e.g. Google Analytics) no longer deliver accurate insight, and when in-platform performance numbers (Meta, TikTok, and other walled gardens) are self-reported, biased, and over-credited. ### Key product capabilities - **Full-funnel marketing reporting** — connect ad spend to real business outcomes - **Cross-channel attribution** — go beyond last-click to measure true ROI across every channel - **Incrementality measurement** — isolate the real causal impact of paid media - **Forecasting & scenario planning** — model optimal spend levels and budget allocation - **Automated budget allocation** — dynamically reallocate spend across channels to maximise ROI ### Advantages - **Independent** — unaffiliated measurement, with no media or platform conflict of interest - **Composable** — integrates with existing systems and runs warehouse-native on the brand's own data (e.g. BigQuery) - **Agentic AI-ready** — the SegmentStream MCP server brings advanced marketing analytics directly into Claude, Codex, and other AI tools - **Transparent** — open, auditable methodology with no "black-box" models ### Core ICP - CMOs, Heads of Performance Marketing, and Digital Marketing Directors - B2C and B2B companies investing $1M–$20M+ annually in digital advertising - Particular strength in lead-generation businesses with longer sales cycles - Clients across the US, UK, and EU ### About the company - Founded in 2018 - VC-backed — raised $3M in seed funding - Headquartered in London, UK - $1–2M ARR range, profitable - Remote team of 5 ## Strategic fit We're selectively open to conversations where the strategic fit is clear and the combined offering beats the sum of its parts. ### Web & Product Analytics - These platforms are strong at on-site and in-product behavior: sessions, funnels, feature adoption, retention, heatmaps. None of them measure the marketing that brought the user in, or the revenue that spend returned. The data stops at the edge of the website. - SegmentStream adds the missing marketing layer on top of behavior analytics: cross-channel attribution, incrementality testing, and automated budget optimization. It connects every session back to the ad, channel, and campaign that created it, and forward to the revenue it produced. The product becomes full-funnel marketing-performance measurement, not just product analytics. - Synergy: the combination closes the loop from the ad that won the customer, to what they did in the product, to the revenue they generated — a story no behavior-analytics platform can tell on its own. It opens a second, larger budget inside every existing account: marketing measurement, owned by the CMO. And because SegmentStream is AI-native infrastructure with an MCP server, that full-funnel dataset becomes something AI agents can query and act on directly, moving the platform from dashboards people read to a measurement layer AI agents run on. ### Digital Agencies / Holdings - Most agencies still report results using the numbers the ad platforms hand them. Those numbers are self-reported, biased, and not independent, and the agency owns no measurement technology of its own. When a client questions the results, the agency has nothing of its own to point to. - SegmentStream gives the agency independent, media-agnostic measurement it owns outright, running warehouse-native on each client's own data. Cross-channel attribution, incrementality testing, and budget recommendations become a product the agency delivers, instead of a third-party fee it pays and marks up. - Synergy: independent, owned measurement becomes a recurring, high-margin product the group rolls out across its entire client base, turning a cost it pays today into proprietary IP it controls. More than that, it is AI-native infrastructure: SegmentStream's MCP server lets the agency's own AI agents run analysis, attribution, and budget decisions across every client automatically, replacing the manual reporting labor that eats agency margins today. Every holding company is racing to become an AI-native operator, and this is the measurement brain that transformation runs on — owned at the group level, sold into accounts they already control, and defensible in a way a media-buying relationship never was. ### Mobile Attribution / MMPs - MMPs are the standard for mobile measurement: app install attribution, in-app events, deep relationships with mobile-first advertisers. Their world is mostly inside the app. Web, cross-channel budgets, and the question of which spend actually drove incremental installs sit largely outside what they measure, and signal loss from SKAdNetwork and privacy changes has made deterministic mobile tracking harder every year. - SegmentStream adds web measurement, cross-channel attribution, incrementality testing, and automated budget optimization on top of the mobile foundation MMPs already own. The result is one measurement layer across app and web, deterministic where the data allows and modeled where it has to be. - Synergy: the combination covers app, web, and every paid channel in one place, exactly as deterministic mobile signal disappears and advertisers demand cross-channel proof. It carries the MMP out of mobile measurement and into the far larger cross-channel budget, where its own customers already spend most of their money. And as media buying shifts to AI agents, those agents need a single cross-channel measurement brain to act on — SegmentStream's warehouse-native, MCP-ready architecture is built to be exactly that, future-proofing the MMP for the agentic, privacy-first era. ### Composable CDPs & Data Integration Platforms - These platforms collect, model, and move customer data: event pipelines, reverse ETL, identity stitching, activation into downstream tools. They run the plumbing well. What they don't tell a customer is whether any of the data they activate actually improved marketing results. - SegmentStream is a measurement application that runs on the same warehouse, using the data these platforms already manage. No new architecture, no extra integration. It turns the raw event and customer data they move into cross-channel attribution, incrementality, and concrete budget decisions. - Synergy: data movement is commoditizing on price and volume, while measurement is the premium outcome customers actually pay up for. Bolting it on gives every customer a direct reason to push more data through the platform, lifting consumption and revenue per account at the same time. It also repositions the company from infrastructure the engineering team manages into a decision tool the CMO opens every week, and through SegmentStream's MCP server, a measurement layer AI agents can query and act on directly, which is exactly where data activation is heading. ### Ad Optimization - These engines automate bidding, budget pacing, and creative decisions across paid media. Every one of those decisions is only as good as the measurement signal feeding it. Today that signal is usually broken: last-click attribution and in-platform, self-reported numbers that over-credit the platforms doing the spending. - SegmentStream provides the accurate measurement the optimization depends on: independent, cross-channel attribution and incrementality testing that show what each channel and campaign actually caused. The engine then makes its bidding and creative calls on reliable data, instead of last-click or in-platform numbers. - Synergy: optimization is only as valuable as the signal it runs on, and today most engines optimize toward broken last-click and in-platform numbers. Putting accurate incrementality and attribution underneath the engine creates a closed loop where measurement, optimization, and activation reinforce each other instead of working against each other. As optimization goes agentic and bidding moves to AI, that loop matters even more: an autonomous optimization agent is only as good as the measurement brain feeding it, and SegmentStream is built to be that brain. It compounds the engine's results and adds an independent measurement product the vendor can sell into the same accounts. ### Fraud Detection & Viewability - These products confirm that media spend is legitimate: real users, viewable placements, brand-safe inventory, no bots or invalid traffic. They verify the quality of the spend. They stop short of saying whether that spend actually grew the business. - SegmentStream measures the next question: did the validated spend produce incremental revenue? It layers cross-channel attribution and incrementality testing on top of the verification and quality signal these companies already own. - Synergy: these vendors already own the trust layer that protects ad budgets. Extending from “the spend was real” to “the spend worked” is the natural next step, sold to the very same buyer, and it carries the company off a defensive, capped loss-prevention budget and onto the far larger performance budget tied to growth. As more spend is allocated by AI agents working from automated signals, owning both the verification and the outcome measurement those agents trust becomes a much stronger position: same customers, same trust, a budget several times the size. ### Marketing Clouds - The big suites are broad across CRM, email, content, and commerce. Cross-channel attribution and marketing measurement are a known soft spot, and the one customers complain about most. Each suite is also under pressure to show a credible, AI-native, warehouse-friendly data story. - SegmentStream drops in the missing measurement and optimization module: warehouse-native, cross-channel, incrementality-backed, and AI-agent-ready out of the box. It runs on the customer's own data instead of forcing everything back into the suite. - Synergy: measurement and cross-channel attribution are the one capability customers ask these suites for again and again, and the one they're most often sent elsewhere to buy. Owning it closes that gap across an installed base measured in the tens of thousands of accounts, turning a churn risk into an upsell. Just as important, SegmentStream arrives as warehouse-native, AI-native infrastructure with an MCP server — exactly the agentic, composable story these suites need to defend against best-of-breed challengers and to make their own AI assistants genuinely useful on marketing data. ### Cloud Data Warehouses - These platforms grow as customers store and process more data, and they increasingly want polished applications that reach business buyers, not only data engineers. Marketing is one of the largest data consumers in any company, and one of the least served by a native application on the platform. - SegmentStream is a ready-made marketing measurement application that runs natively on the warehouse, on Snowflake, BigQuery, or Databricks. It drives real query and compute usage, and it is sold to the marketing team rather than the data team. - Synergy: a marketing-measurement workload is one of the highest-volume, most recurring compute workloads a marketing org can run, which is exactly the consumption these platforms monetize. The acquisition also opens a direct line to the CMO and the marketing budget, a buyer these platforms rarely reach without a partner. And as AI agents start running real work on the warehouse, SegmentStream's MCP server makes marketing one of the first agent-driven workloads on the platform — a flagship, AI-native application that drives consumption and proves the platform's agentic story at the same time. ### Customer Engagement & ESPs - These platforms run lifecycle and retention marketing: profiles, segmentation, email, push, and onboarding flows. They optimize value after the customer is already in the door. They see very little of what it cost to acquire that customer, or which acquisition spend actually brings in the profitable ones. - SegmentStream's attribution links acquisition to lifetime-value outcomes. Teams already using these platforms to maximise customer lifetime value can now also see which ad channels and activities bring in the customers who end up most valuable — 360° analysis and optimization across the full funnel, not just what happens after a customer is acquired. - Synergy: these platforms own everything that happens after acquisition; SegmentStream owns the spend that drives it. Together they optimize the entire customer lifecycle in one place, from the ad that wins a customer to the lifetime value they go on to generate. That connection opens the acquisition budget, typically far larger than the retention budget these tools serve today. And it gives the AI agents these platforms are racing to ship a full-funnel signal to reason over, instead of the retention-only picture they have today. ### Marketing Mix Modeling (MMM) - Marketing mix modeling estimates channel effectiveness from the top down, in periodic and fairly broad studies. Buyers increasingly want those models confirmed with always-on, granular, experiment-based measurement, and the two methods are most convincing when they agree. - SegmentStream adds actionable, granular measurement and optimization through deterministic attribution and incrementality testing — so CMOs can move beyond strategic, annual MMM reports to continuous, always-on optimization. They see which exact creative or campaign actually worked, not just the Offline vs TV vs Radio vs Digital mix. - Synergy: owning both methods lets the vendor validate the model against always-on, experiment-based reality, replace slow annual refreshes with measurement that updates every week, and sell one complete measurement system instead of a periodic study. Top-down MMM sets the strategy; bottom-up attribution and incrementality run the day-to-day. And as planning and optimization move to AI agents that act continuously, periodic studies stop being enough on their own: the combined product gives those agents the always-on, granular signal they need, which is exactly where the measurement market is consolidating. ### AI / LLM Search Optimization - This young category tracks brand visibility and citations inside AI assistants like ChatGPT, Gemini, and Perplexity, often called generative engine optimization or answer engine optimization. It can measure presence and share of voice. It can't yet tell a brand what that presence is actually worth in revenue. - SegmentStream's measurement technology captures the true incremental impact and ROI of AI search on final business outcomes — often 5–10x what traditional analytics tools like Google Analytics can track by observing referral traffic with last-click attribution. - Synergy: visibility metrics are hard to monetize on their own, because no CFO funds share of voice without a line to revenue. Outcome measurement turns AI-search presence into a hard ROI number, which is what unlocks real budget for the category. Both products are AI-native by design and share an MCP-ready architecture, so they snap together, and the combination sits early in one of the fastest-growing, most AI-native corners of measurement, built for a world where AI agents both create and consume the search. ## Integration readiness The risks acquirers price in — answered by architecture. Every small-team acquisition gets discounted for the same four risks. Here they are, with the answers technical diligence will find — in the architecture, not in assurances. - **Integration risk — Will this take two years to integrate?** No — the platform is already packaged for embedding: a standalone managed cloud with workspace-per-client isolation, EU & US regions, and REST / TypeScript SDK / MCP surfaces. Integrating means an event-source adapter and an SDK embed, not a replatforming program. - **Key-person risk — Is the value locked in five people's heads?** The engine runs on a newly rewritten codebase that deploys as a self-contained cloud platform — documented, region-portable, and operable by any competent platform team or by AI agents. The codebase is structured agent-first, with per-module agent instructions, so it can be developed and extended entirely by agents. The asset is the platform, not tribal knowledge. - **Technology diligence — Is the embeddable story real, or a slide?** It's a live product surface, not a roadmap. SegmentStream Cloud offers this same engine, white-label, today — CLI-provisioned workspaces, an embeddable measurement module, and metered usage. See [SegmentStream Cloud](https://segmentstream.com/cloud.md). - **Customer entanglement — Will clients churn through a migration?** The product is warehouse-native by design: customer data lives in the customer's own warehouse, one workspace per client. There is no proprietary data store to migrate off — continuity through an acquisition is the default, not a project. ## Founders - **Constantine Yurevich** — Co-Founder · [LinkedIn](https://www.linkedin.com/in/yurevichcv/) · [X](https://x.com/weird_ceo) - **Pavel Petrinich** — Co-Founder · [LinkedIn](https://www.linkedin.com/in/pavelpetrinich/) · [X](https://x.com/pavel_petrinich) ## Frequently asked questions **What does SegmentStream do?** SegmentStream is an independent, AI-native marketing measurement platform. It connects ad spend to real business outcomes with cross-channel attribution, incrementality testing, marketing mix optimization, and automated budget allocation, running warehouse-native on the brand's own data. The same engine is available to AI agents through an MCP server, so teams and their AI tools work from one source of measurement truth. **What market category is SegmentStream in?** Marketing measurement and marketing analytics: cross-channel attribution, incrementality testing, marketing mix optimization, paid media intelligence, and agentic AI ad optimization. It competes in the same market as attribution and ad-measurement vendors, built AI-native from the ground up. **What makes SegmentStream different, and what is defensible?** Four things. It is independent, with no media or platform conflict of interest. It is composable and warehouse-native, running on the customer's own data in BigQuery, Snowflake, or Databricks, with no data lock-in. It is AI-agent-ready through its MCP server. And it is transparent, with open, auditable methodology and no black-box models. The defensible core is the measurement engine itself: years of work modeling attribution and incrementality across 30+ ad platforms, now packaged as infrastructure AI agents can build on. **How does SegmentStream fit the shift to AI and agents?** SegmentStream is built as infrastructure that gives AI agents a marketing measurement brain. Its MCP server brings full-funnel attribution, incrementality, and budget optimization directly into Claude, Cursor, Codex, and any MCP client, so an agent can ask a measurement question and act on the answer. As marketing moves from dashboards people read to agents that plan and optimize continuously, that measurement layer is what those agents run on. **Who uses SegmentStream?** Brands across e-commerce, B2B SaaS, finance, travel, energy, and automotive, typically investing $1M–$20M+ a year in digital advertising, with particular strength in lead-generation businesses that have longer sales cycles. The measurement engine is used across more than $100M in annual ad spend, and SegmentStream is rated 4.7/5 on G2. **Is SegmentStream profitable, and how large is it?** Yes. SegmentStream is profitable, in the $1–2M ARR range, with a lean remote team of five. It was founded in 2018, raised $3M in seed funding, and is headquartered in London, UK, serving clients across the US, UK, and EU. **Is SegmentStream open to acquisition?** Yes, selectively. SegmentStream is open to acquisition proposals from strategic buyers positioned to scale the technology, team, and category position, where the combined offering beats the sum of its parts. Inquiries on this page go directly to the founders. **Which kinds of companies are the strongest strategic fit to acquire SegmentStream?** Web and product analytics platforms, agencies and holding companies, mobile attribution providers (MMPs), CDPs and data-integration platforms, ad-optimization engines, customer engagement platforms and ESPs, marketing clouds, cloud data warehouses, MMM vendors, and AI/LLM search-optimization tools. The strategic-fit section above explains why each is a fit. --- # Cut Through the Noise — Master Marketing Measurement That Actually Reflects Reality On-demand course by Constantine Yurevich. 9.5 hours of no-BS video lessons on attribution, MMM, and incrementality from first principles. Learn to see through platform bias and make confident, evidence-based marketing decisions. ## Why this course exists Marketing measurement has become a belief system. Most teams trust attribution models and MMM frameworks they can’t explain. This course is for marketers who want evidence — not faith. It rebuilds the discipline from first principles — so you can challenge models, run real experiments, and make budget decisions you can actually defend. ## Stats - 9.5h video lessons - 4 modules - 52 topics - $299 lifetime access (one-time payment) ## Modules ### 01. Attribution and re-attribution (2.5h, 15 topics) Why we need analytics. Last-click attribution. Ad platform (post-click) attribution. Post-click + post-view attribution. Meta "incremental attribution". "Cookieless" attribution. Multi-touch and first-click attribution. Identity graph. Click-time reporting. Conversion modeling for conversion maturation. Noise smoothing. Conversion modeling to recover consent losses. Re-attribution methodology. ### 02. Optimization and budget allocation (2.5h, 15 topics) How to make attribution insights actionable. Marginal analytics. Practical example: where would you allocate additional $20,000? Ad platforms' target ROAS. Upper / mid / lower funnel. Splitting budget across channels and campaigns. In-platform optimization. Hidden costs of splitting campaigns. Conversion API to maximize tracked conversions. What to do when you lack directly attributed conversion. Common mistakes when testing new signals and bidding strategies. A/B test vs new campaign in parallel. Ad experiments statistical significance calculator. ### 03. Incrementality testing and MMM (3h, 14 topics) Incrementality testing. In-platform lift test. Geo holdout test. Minimum detectable effect. Common mistakes when running geo holdout test. Placebo test. Use cases for geo holdout test. What is marketing mix modeling? Main difficulties of MMM. Real MMM story. Next-gen MMM. How priors can overcredit ad platform's impact. "Causal" MMM. Triangulation. ### 04. Lead gen and LTV-based measurement (1.5h, 8 topics) Main challenge of lead gen business. Lead quality: how to properly link online and offline data. Lead scoring. Example how it works. Must-haves for B2B measurement. LTV-focused business. Key mistakes when analysing LTV. Customer LTV prediction. ## Instructor Constantine Yurevich — Founder of SegmentStream. A decade helping advanced marketing teams move from biased measurement methodologies to incrementality- and evidence-first approaches. The course distils SegmentStream's practical methods into nine and a half hours of no-fluff instruction. ## Who it's for Marketing teams running $10k+/month in ad spend. Agencies, marketing leaders, founders, and digital marketing managers who need to make defensible budget decisions across channels. ## What's included - 4 video modules · 9.5 hours of lessons - Lifetime access to all current and future updates - Certificate of Theoretical Completion (shareable on LinkedIn, verifiable, unique ID, 10 training points) - Self-paced — start anytime ## Pricing $299 one-time payment. Lifetime access. All future updates. Secure payment via Stripe. Instant access on successful payment. Non-refundable due to instant digital delivery. Purchase: https://course.segmentstream.com/purchase ## FAQ - **Access**: Lifetime access to all course materials and future updates. - **Pace**: Fully on-demand. Start any time, move at your own speed. - **Format**: Four video modules totaling 9.5 hours, plus downloadable resources including calculators and reference materials. Recorded from the live cohort led by Constantine in October 2025. - **Prerequisites**: No technical background required. Designed for both beginners and experienced marketers. - **Refunds**: Non-refundable due to instant digital access. Standard for digital information products. - **Certificate**: Complete all four modules to receive a Certificate of Theoretical Completion — shareable on LinkedIn with a unique verification ID and 10 training points. - **Outcomes**: Complete understanding of modern marketing measurement — moving beyond belief-based methodologies to evidence-based, statistically valid approaches. ## Featured podcast appearances - DTC Podcast — The hard truth about marketing measurement (61 min) - Chew On This Podcast — What DTC brands get wrong about attribution (52 min) - DTC Podcast — Marginal ROAS and how Meta's algorithm really works (49 min) --- # Trust & Security ## How we handle your customer data SegmentStream offers two data storage options. You choose during setup: **Self-hosted storage** — you provide your own Google BigQuery project. SegmentStream connects to it, reads from it, and writes analytics results back to it. We do not maintain a separate copy of your marketing data. You retain full control of the infrastructure. **SegmentStream-hosted storage** — we provide a managed data warehouse for you. Your data is stored in an isolated dataset on Google Cloud Platform in the region you select (EU or US). You retain full ownership of your data and can export or migrate to your own infrastructure at any time. Regardless of which option you choose: - **Depending on the context, we act as a data processor or a data controller.** We are a data processor when processing data on your behalf within the Platform, and a data controller for our Website, marketing, and business operations. Under GDPR and equivalent frameworks, Platform processing is carried out on your behalf and under your instructions. Our Data Processing Agreement formalises this relationship. - **Your data is isolated.** Each customer's data is stored in a separate dataset. There is no commingling of data between customers. - **Ad platform credentials are used for read-only data collection.** When you connect ad platforms (Google Ads, Meta, etc.), we collect cost and performance data and load it into your warehouse. We do not modify your ad accounts. - **Data portability.** You can migrate from SegmentStream-hosted storage to your own infrastructure at any time. We support data transfer to your own Google BigQuery project. - **Deletion.** If you request account deletion, all data in SegmentStream-hosted storage is removed from active systems promptly. Complete purge from all underlying storage (including Google Cloud's built-in recovery mechanisms) occurs within 14 days. - **Sub-processors are disclosed.** A full list of sub-processors is available in Annex D of our [Data Processing Agreement](https://segmentstream.com/dpa.md). ## Compliance Our legal and data protection framework covers the major regulatory regimes: - **GDPR** — EU General Data Protection Regulation - **UK GDPR** — UK Data Protection Act 2018 - **CCPA** — California Consumer Privacy Act - **PIPEDA** — Canadian Personal Information Protection and Electronic Documents Act - **LGPD** — Brazilian General Data Protection Law - **Swiss DPA** — Swiss Federal Data Protection Act Standard Contractual Clauses (SCCs) are included in our DPA for international data transfers where required. ## Security - Automated security monitoring and compliance tracking via [Drata](https://app.drata.com/trust/9cb9666c-0c38-11ee-865f-029d78a187d9) - Encryption in transit (TLS) and at rest - Role-based access controls with least-privilege principles - Regular access reviews and audit logging - Infrastructure hosted on Google Cloud Platform ## How we use platform interaction data SegmentStream uses an AI agent that connects to your data via the MCP protocol to query and analyse it in real time. When you or your team interact with the agent — asking questions, running reports, connecting data sources — we log those interactions (prompts, tool calls, agent responses, and session metadata). We use this data for two purposes: - **Debugging and support.** When you report an issue — for example, the agent set up a project incorrectly or used a tool in an unexpected way — our team investigates the specific session logs to diagnose the problem and resolve it. - **Product improvement.** We review interaction patterns and logs in a controlled and access-restricted manner to improve system performance and reliability. This does not involve training general-purpose AI models on your data. What we do not do: - We do not use your data, including prompts, interaction logs, or Customer Data, to train general-purpose AI models. - We do not use your marketing data (the data in your BigQuery) to train AI models or improve our product. Your warehouse data is processed solely to answer your questions and run your reports. - We do not share your interaction logs with other customers. - We do not sell any data. This is documented in Sections 3.2.4 and 4.2.4 of our [Privacy Policy](https://segmentstream.com/privacy.md). ## Confidentiality Our [Terms of Service](https://segmentstream.com/terms.md) include mutual confidentiality obligations (Section 8). This means: - All information you share with us — including Customer Data, business plans, and technical details — is protected under contractual confidentiality. - We will not disclose your confidential information to third parties except to employees and contractors who need to know and are bound by equivalent obligations. - These obligations survive termination for three years. In most cases, this removes the need for a separate NDA, although we can support additional agreements where required. Custom legal agreements are available for enterprise paid plans. ## Legal documents - [Terms of Service](https://segmentstream.com/terms.md) — your agreement with SegmentStream, including confidentiality obligations - [Privacy Policy](https://segmentstream.com/privacy.md) — how we collect and use personal data - [Data Processing Agreement](https://segmentstream.com/dpa.md) — GDPR-compliant DPA with SCCs, sub-processor list, and security measures - [Cookie Policy](https://segmentstream.com/cookies.md) — cookies used on this website ## Frequently asked questions ### Do you sign NDAs? Our Terms of Service include mutual confidentiality obligations that cover all information exchanged between us. For most use cases — including Research Preview — this provides equivalent protection to a standalone NDA. Custom legal agreements are available for enterprise paid plans. ### Where is my data stored? If you use self-hosted storage, your marketing data stays in your own BigQuery project. If you use SegmentStream-hosted storage, your data is stored in an isolated dataset on Google Cloud Platform in the region you selected (EU or US). In both cases, SegmentStream's application infrastructure runs on Google Cloud Platform. For details on data processing locations and sub-processors, see Annex C and Annex D of our [Data Processing Agreement](https://segmentstream.com/dpa.md). ### Can I migrate from hosted to self-hosted storage? Yes. You can migrate your data from SegmentStream-hosted storage to your own Google BigQuery project at any time. Contact us and we will arrange the transfer. ### What happens to my data if I cancel? For self-hosted storage, your data remains in your own BigQuery project — we simply disconnect. For SegmentStream-hosted storage, we will provide a window to export your data, after which it is deleted from our systems. Data is removed from active systems promptly upon deletion, with complete purge from all underlying storage within 14 days. ### Who are your sub-processors? A full list is maintained in Annex D of our [Data Processing Agreement](https://segmentstream.com/dpa.md), including each sub-processor's purpose, location, and the data they process. ### Do you have SOC 2? We maintain continuous security and compliance monitoring through Drata. Our security controls are aligned with SOC 2 principles. Our current security posture is available at our [Drata Trust Center](https://app.drata.com/trust/9cb9666c-0c38-11ee-865f-029d78a187d9). ### What should I send to my legal or procurement team? Share this page along with our [Terms of Service](https://segmentstream.com/terms.md) and [Data Processing Agreement](https://segmentstream.com/dpa.md). These documents cover confidentiality, data processing, security measures, sub-processors, and international transfer mechanisms — everything a legal or procurement review typically requires. ### Is the Research Preview subject to the same protections? Yes. All Customer Data processed during the Research Preview is subject to our Data Processing Agreement and the same security, confidentiality, and data protection obligations that apply to our paid services. This is stated explicitly in Section 7 of our [Terms of Service](https://segmentstream.com/terms.md). --- # Google Analytics 4 Integration How SegmentStream integrates with Google Analytics 4: reads raw event data from the GA4 → BigQuery export to power cross-channel attribution and budget optimization. SegmentStream layers cross-channel attribution and budget optimization on top of your existing Google Analytics 4 tracking — using raw event data from the GA4 → BigQuery export, not the GA4 Reporting API. ## How the integration works SegmentStream reads from your [Google Analytics 4 → BigQuery export](https://support.google.com/analytics/answer/9358801) on a daily schedule. SegmentStream consumes the daily (`events_YYYYMMDD`) tables, and falls back to intraday (`events_intraday_YYYYMMDD`) when the daily batch hasn't landed yet — so workflows don't block waiting on it. ## Raw events only — not the GA4 Reporting API SegmentStream does **not** use the [GA4 Reporting API](https://developers.google.com/analytics/devguides/reporting/data/v1). The API exposes only processed data, with GA4's own session definitions and attribution model baked in. SegmentStream consumes the raw exported events and builds attribution using its own identity graph and algorithms — which is the entire point of layering on top of GA4 instead of just re-reading what GA4 reports. ## Conversions export to GA4 SegmentStream can be configured to export conversions to Google Analytics 4 — including events that don't originate on the website. Any conversion SegmentStream processes from an external data source can be exported to GA4 and tied to the original GA4 user / session, so reports reflect real outcomes instead of just on-site events. This is a capability, not on by default — you choose which conversions to export. Examples of what you can export: - **CRM / sales-pipeline events** — deeper-funnel signals like qualified-lead, opportunity-stage transitions, and closed-won, so paid traffic that drove pipeline (not just form fills) shows up in GA4 reporting. - **E-commerce post-purchase events** — refunds, returns, and shipped orders, so GA4 conversion data reflects real revenue net of returns instead of just initial checkouts. - **Any offline event** — phone-call conversions, in-store sales, or anything SegmentStream ingests from CRM / commerce / call-tracking sources. ## Historical data Historical data is only available as far back as the BigQuery export was enabled. GA4 does not retro-fill past data when the link is created, and [once exported it cannot be re-exported](https://support.google.com/analytics/answer/9358801). Recommendation: enable the BigQuery export early — even before signing — so that the history is already there by the time you start using SegmentStream. ## At a Glance - **Read**: GA4 → BigQuery export (raw events) - **Where the data lives**: Your BigQuery - **Latency**: Daily - **Historical data**: From the date the BigQuery export was enabled - **Export to GA4**: Conversions — CRM, e-commerce, offline events ## Official Links - [GA4 BigQuery export setup (Google support)](https://support.google.com/analytics/answer/9358801) - [GA4 → BigQuery schema reference](https://support.google.com/analytics/answer/7029846) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Adobe Analytics Integration How SegmentStream integrates with Adobe Analytics: reads raw hit-level data from Adobe Data Feeds delivered to a SegmentStream-provisioned Google Cloud Storage bucket, then layers attribution and budget optimization on top. SegmentStream layers cross-channel attribution and budget optimization on top of your existing Adobe Analytics tracking — using raw hit-level data from Adobe Data Feeds, not the Adobe Analytics Reporting API. ## How the integration works SegmentStream reads from [Adobe Analytics Data Feeds](https://experienceleague.adobe.com/en/docs/analytics/export/analytics-data-feed/data-feed-overview) on a daily schedule. SegmentStream provisions a Google Cloud Storage bucket and reads each daily feed from there. Setup: in Adobe Analytics, configure a [Google Cloud Platform export location](https://experienceleague.adobe.com/en/docs/analytics/components/locations/configure-import-locations). Adobe issues a **Principal** (the identity Adobe writes from). You hand that Principal to SegmentStream — SegmentStream provisions the bucket and grants the Principal write-access automatically. From then on, daily feeds land in the bucket and SegmentStream picks them up. Data retention: SegmentStream removes feed files from the GCS bucket **seven days after they're processed**. The SegmentStream-processed data lives in your warehouse alongside the rest of your SegmentStream datasets. ## Raw events only — not the Adobe Analytics Reporting API SegmentStream does **not** use the Adobe Analytics Reporting API. The API exposes only processed reports with Adobe's own visit / visitor definitions and attribution. SegmentStream consumes the raw exported hits (every event, every eVar / prop / event variable) and builds attribution using its own identity graph and algorithms — that's the entire point of layering on top of Adobe instead of re-reading what Adobe reports. ## Historical data When creating an Adobe Data Feed you can set the [feed start date to any date in the past where data is being collected](https://experienceleague.adobe.com/en/docs/analytics/export/analytics-data-feed/create-feed), and Adobe will replay historical hits into the destination bucket. The only practical limit is that the report suite must have been actively collecting data for the period you want. ## At a Glance - **Read**: Adobe Data Feeds (raw events) - **Where the data lives**: SegmentStream-managed GCS bucket; processed data in your warehouse - **Latency**: Daily - **Historical data**: Replay any past period the report suite was collecting ## Official Links - [Adobe Data Feeds overview](https://experienceleague.adobe.com/en/docs/analytics/export/analytics-data-feed/data-feed-overview) - [Create an Adobe Data Feed](https://experienceleague.adobe.com/en/docs/analytics/export/analytics-data-feed/create-feed) - [Configure an Adobe export location (GCP, Principal setup)](https://experienceleague.adobe.com/en/docs/analytics/components/locations/configure-import-locations) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Heap by Contentsquare Integration How SegmentStream integrates with Heap by Contentsquare: reads raw event data from Heap Connect — exported to your data warehouse — and layers identity graph and cross-channel attribution on top. SegmentStream layers cross-channel attribution and budget optimization on top of Heap by Contentsquare — reading raw event data exported from Heap Connect into your data warehouse, not the Heap reporting interface. ## How the integration works Heap Connect exports raw event data into your own data warehouse. Once that export is configured, point SegmentStream at the same warehouse and SegmentStream reads the events from there. [Heap Connect](https://docs.contentsquare.com/en/connect/data-warehouses-overview/) supports [BigQuery](https://docs.contentsquare.com/en/connect/bigquery/), [Snowflake](https://docs.contentsquare.com/en/connect/snowflake/), [Redshift](https://docs.contentsquare.com/en/connect/redshift/), [Databricks](https://docs.contentsquare.com/en/connect/databricks/), and [Amazon S3](https://docs.contentsquare.com/en/connect/amazon-s3/). SegmentStream connects to whichever one you use. ## Raw events only — not the Heap Reporting API SegmentStream does **not** use Heap's reporting API. The reporting API exposes only processed data with Heap's own session and identity model. SegmentStream consumes the raw exported events from Heap Connect and builds attribution using its own identity graph and algorithms — the same principle as for every other behavioral analytics source. ## Historical data On initial connection, Heap Connect creates the dataset in your warehouse and populates it with available historical event data. SegmentStream becomes able to use that history as soon as the dataset is reachable. For exact retention and historical-coverage limits on your Heap plan, see Heap's [Data Connect documentation](https://docs.contentsquare.com/en/connect/data-warehouses-overview/). ## At a Glance - **Read**: Heap Connect (raw events) - **Where the data lives**: Your warehouse — [BigQuery](https://docs.contentsquare.com/en/connect/bigquery/), [Snowflake](https://docs.contentsquare.com/en/connect/snowflake/), [Redshift](https://docs.contentsquare.com/en/connect/redshift/), [Databricks](https://docs.contentsquare.com/en/connect/databricks/), or [Amazon S3](https://docs.contentsquare.com/en/connect/amazon-s3/) - **Latency**: Daily - **Historical data**: Heap populates the dataset on initial sync ## Official Links - [Heap Connect — supported data warehouses (Contentsquare docs)](https://docs.contentsquare.com/en/connect/data-warehouses-overview/) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Amplitude Integration How SegmentStream integrates with Amplitude: reads raw event data exported from Amplitude into your data warehouse, then layers identity graph and cross-channel attribution on top. SegmentStream layers cross-channel attribution and budget optimization on top of Amplitude — reading raw event data exported from Amplitude's Destination Catalog into your data warehouse, not the Amplitude reporting interface. ## How the integration works Amplitude exports raw event data into your own data warehouse via its warehouse destinations. Once that export is configured, point SegmentStream at the same warehouse and SegmentStream reads the events from there. [Amplitude](https://amplitude.com/docs/data/destination-catalog) supports [BigQuery](https://amplitude.com/docs/data/destination-catalog/google-bigquery), [Snowflake](https://amplitude.com/docs/data/destination-catalog/snowflake) (and [Data Share](https://amplitude.com/docs/data/destination-catalog/snowflake-data-share)), [Redshift](https://amplitude.com/docs/data/destination-catalog/amazon-redshift), [Amazon S3](https://amplitude.com/docs/data/destination-catalog/amazon-s3), [GCS](https://amplitude.com/docs/data/destination-catalog/google-cloud-storage), and [Azure Blob Storage](https://amplitude.com/docs/data/destination-catalog/azure-blob-storage). SegmentStream connects to whichever one you use. ## Raw events only — not the Amplitude Reporting API SegmentStream does **not** use Amplitude's reporting / dashboards API. Those expose only processed data with Amplitude's own session and identity model. SegmentStream consumes the raw exported events and builds attribution using its own identity graph and algorithms. ## Historical data On initial setup, Amplitude can export historical events alongside the forward feed. SegmentStream becomes able to use that history as soon as the warehouse is reachable. For exact retention and historical-coverage limits, see Amplitude's [Destination Catalog documentation](https://amplitude.com/docs/data/destination-catalog). ## At a Glance - **Read**: Amplitude warehouse destinations (raw events) - **Where the data lives**: Your warehouse or object store — [BigQuery](https://amplitude.com/docs/data/destination-catalog/google-bigquery), [Snowflake](https://amplitude.com/docs/data/destination-catalog/snowflake), [Redshift](https://amplitude.com/docs/data/destination-catalog/amazon-redshift), [Amazon S3](https://amplitude.com/docs/data/destination-catalog/amazon-s3), [GCS](https://amplitude.com/docs/data/destination-catalog/google-cloud-storage), or [Azure Blob](https://amplitude.com/docs/data/destination-catalog/azure-blob-storage) - **Latency**: Daily - **Historical data**: Amplitude exports historical events on initial setup ## Official Links - [Amplitude Destination Catalog](https://amplitude.com/docs/data/destination-catalog) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Mixpanel Integration How SegmentStream integrates with Mixpanel: reads raw event data exported via Mixpanel Data Pipelines into your data warehouse, layered with identity graph and cross-channel attribution. SegmentStream layers cross-channel attribution and budget optimization on top of Mixpanel — reading raw event data exported by Mixpanel Data Pipelines into your data warehouse, not the Mixpanel reporting interface. ## How the integration works Mixpanel Data Pipelines export raw event data, user profiles, and identity mappings into your own data warehouse on a continuous schedule. Once that export is configured, point SegmentStream at the same warehouse and SegmentStream reads the events from there. [Mixpanel Data Pipelines](https://docs.mixpanel.com/docs/data-pipelines/integrations) support [BigQuery](https://docs.mixpanel.com/docs/data-pipelines/integrations/bigquery), [Snowflake](https://docs.mixpanel.com/docs/data-pipelines/integrations/snowflake), [Databricks](https://docs.mixpanel.com/docs/data-pipelines/integrations/databricks), [Redshift Spectrum](https://docs.mixpanel.com/docs/data-pipelines/integrations/redshift-spectrum), [AWS S3](https://docs.mixpanel.com/docs/data-pipelines/integrations/aws-s3), [GCS](https://docs.mixpanel.com/docs/data-pipelines/integrations/gcp-gcs), and [Azure Blob Storage](https://docs.mixpanel.com/docs/data-pipelines/integrations/azure-blob-storage). SegmentStream connects to whichever one you use. ## Raw events only — not the Mixpanel Reporting API SegmentStream does **not** use Mixpanel's query / reporting API. That API exposes only processed data with Mixpanel's own session and identity model. SegmentStream consumes the raw exported events and builds attribution using its own identity graph and algorithms. ## Historical data On initial connection, Mixpanel can export historical events into the warehouse alongside the ongoing pipeline. SegmentStream becomes able to use that history as soon as the warehouse is reachable. Data Pipelines is an Enterprise / Growth add-on; see Mixpanel's [Data Pipelines documentation](https://docs.mixpanel.com/docs/data-pipelines) for retention and access details. ## At a Glance - **Read**: Mixpanel Data Pipelines (raw events) - **Where the data lives**: Your warehouse or object store — [BigQuery](https://docs.mixpanel.com/docs/data-pipelines/integrations/bigquery), [Snowflake](https://docs.mixpanel.com/docs/data-pipelines/integrations/snowflake), [Databricks](https://docs.mixpanel.com/docs/data-pipelines/integrations/databricks), [Redshift Spectrum](https://docs.mixpanel.com/docs/data-pipelines/integrations/redshift-spectrum), [AWS S3](https://docs.mixpanel.com/docs/data-pipelines/integrations/aws-s3), [GCS](https://docs.mixpanel.com/docs/data-pipelines/integrations/gcp-gcs), or [Azure Blob](https://docs.mixpanel.com/docs/data-pipelines/integrations/azure-blob-storage) - **Latency**: Daily - **Historical data**: Mixpanel exports historical events on initial setup ## Official Links - [Mixpanel Data Pipelines overview](https://docs.mixpanel.com/docs/data-pipelines) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # PostHog Integration How SegmentStream integrates with PostHog: reads raw event data via PostHog Batch Exports into your data warehouse, layered with identity graph and cross-channel attribution. SegmentStream layers cross-channel attribution and budget optimization on top of PostHog — reading raw event data exported by PostHog Batch Exports into your data warehouse, not the PostHog reporting interface. ## How the integration works PostHog exports raw event data into your own data warehouse via Batch Exports. Once that export is configured, point SegmentStream at the same warehouse and SegmentStream reads the events from there. [PostHog Batch Exports](https://posthog.com/docs/cdp/batch-exports) support [BigQuery](https://posthog.com/docs/cdp/batch-exports/bigquery), [Snowflake](https://posthog.com/docs/cdp/batch-exports/snowflake), [Redshift](https://posthog.com/docs/cdp/batch-exports/redshift), [Databricks](https://posthog.com/docs/cdp/batch-exports/databricks), [Postgres](https://posthog.com/docs/cdp/batch-exports/postgres), [Amazon S3](https://posthog.com/docs/cdp/batch-exports/s3), and [Azure Blob Storage](https://posthog.com/docs/cdp/batch-exports/azureblob). SegmentStream connects to whichever one you use. ## Raw events only — not the PostHog Reporting API SegmentStream does **not** use PostHog's query / reporting API. That API exposes only processed data with PostHog's own session and identity model. SegmentStream consumes the raw exported events and builds attribution using its own identity graph and algorithms. ## Historical data PostHog Batch Exports run forward from when they're configured. Historical coverage and retention depend on your PostHog plan and instance — see the [Batch Exports documentation](https://posthog.com/docs/cdp/batch-exports) for limits. ## At a Glance - **Read**: PostHog Batch Exports (raw events) - **Where the data lives**: Your warehouse — [BigQuery](https://posthog.com/docs/cdp/batch-exports/bigquery), [Snowflake](https://posthog.com/docs/cdp/batch-exports/snowflake), [Redshift](https://posthog.com/docs/cdp/batch-exports/redshift), [Databricks](https://posthog.com/docs/cdp/batch-exports/databricks), [Postgres](https://posthog.com/docs/cdp/batch-exports/postgres), [Amazon S3](https://posthog.com/docs/cdp/batch-exports/s3), or [Azure Blob](https://posthog.com/docs/cdp/batch-exports/azureblob) - **Latency**: Daily - **Historical data**: Per PostHog plan; exports run forward from setup ## Official Links - [PostHog Batch Exports — supported destinations](https://posthog.com/docs/cdp/batch-exports) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Twilio Segment Integration How SegmentStream integrates with Twilio Segment: reads raw event data routed by Segment into your data warehouse, layered with identity graph and cross-channel attribution. SegmentStream layers cross-channel attribution and budget optimization on top of Twilio Segment — reading raw event data routed by Segment's Warehouse destinations into your data warehouse. ## How the integration works Twilio Segment is a customer data platform — it collects events from your sites and apps and routes them to destinations. One class of destinations is data warehouses. When Segment is configured to write to a warehouse, SegmentStream reads from that same warehouse. [Segment Warehouse destinations](https://www.twilio.com/docs/segment/connections/storage/warehouses) include [BigQuery](https://www.twilio.com/docs/segment/connections/storage/catalog/bigquery), [Snowflake](https://www.twilio.com/docs/segment/connections/storage/catalog/snowflake), [Redshift](https://www.twilio.com/docs/segment/connections/storage/catalog/redshift), [Databricks](https://www.twilio.com/docs/segment/connections/storage/catalog/databricks), [Postgres](https://www.twilio.com/docs/segment/connections/storage/catalog/postgres), [Azure SQL DW](https://www.twilio.com/docs/segment/connections/storage/catalog/azuresqldw), and [Data Lakes (AWS Glue / S3)](https://www.twilio.com/docs/segment/connections/storage/catalog/data-lakes). SegmentStream connects to whichever one you use. ## Segment is the pipe — SegmentStream is the measurement layer Segment routes raw events from your sources to destinations. SegmentStream isn't a Segment destination in the classic sense — it consumes the same raw events that Segment lands in your warehouse, then builds identity graph, cross-channel attribution, and budget optimization on top. No separate event mapping or transformation is required. ## Historical data Segment writes warehouse data forward from when the destination is enabled. For replays of historical event data, see Segment's [Warehouse destinations documentation](https://www.twilio.com/docs/segment/connections/storage/warehouses) — replay support varies by source and plan. ## At a Glance - **Read**: Segment Warehouse destinations (raw events) - **Where the data lives**: Your warehouse or data lake — [BigQuery](https://www.twilio.com/docs/segment/connections/storage/catalog/bigquery), [Snowflake](https://www.twilio.com/docs/segment/connections/storage/catalog/snowflake), [Redshift](https://www.twilio.com/docs/segment/connections/storage/catalog/redshift), [Databricks](https://www.twilio.com/docs/segment/connections/storage/catalog/databricks), [Postgres](https://www.twilio.com/docs/segment/connections/storage/catalog/postgres), [Azure SQL DW](https://www.twilio.com/docs/segment/connections/storage/catalog/azuresqldw), or [AWS Glue / S3 Data Lakes](https://www.twilio.com/docs/segment/connections/storage/catalog/data-lakes) - **Latency**: Daily - **Historical data**: Forward from when the warehouse destination is enabled ## Official Links - [Segment Warehouse destinations](https://www.twilio.com/docs/segment/connections/storage/warehouses) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Snowplow Integration How SegmentStream integrates with Snowplow: reads raw enriched event data loaded by Snowplow into your data warehouse, layered with identity graph and cross-channel attribution. SegmentStream layers cross-channel attribution and budget optimization on top of Snowplow — reading raw enriched event data loaded by Snowplow's warehouse loaders into your data warehouse. ## How the integration works Snowplow collects, enriches, and loads behavioral event data into your own data warehouse. Once Snowplow is loading data, point SegmentStream at the same warehouse and SegmentStream reads the events from there. [Snowplow](https://docs.snowplow.io/docs/fundamentals/destinations/) supports [BigQuery](https://docs.snowplow.io/docs/destinations/warehouses-lakes/bigquery/), [Snowflake](https://docs.snowplow.io/docs/destinations/warehouses-lakes/snowflake/), [Redshift](https://docs.snowplow.io/docs/destinations/warehouses-lakes/redshift/), [Databricks](https://docs.snowplow.io/docs/destinations/warehouses-lakes/databricks/), and the open lake formats [Delta Lake](https://docs.snowplow.io/docs/destinations/warehouses-lakes/delta/) and [Apache Iceberg](https://docs.snowplow.io/docs/destinations/warehouses-lakes/iceberg/). SegmentStream connects to whichever one you use. ## Snowplow is the pipe — SegmentStream is the measurement layer Snowplow's job is to collect and land high-quality, enriched behavioral data in your warehouse. SegmentStream then builds identity graph, cross-channel attribution, and budget optimization on top of those events. Snowplow's schema is read natively — no separate event mapping or transformation is required. ## Historical data Snowplow loads data forward from when the pipeline starts running, and any enrichment changes can be re-applied. Historical replay of raw collector data depends on how your pipeline retains it — see Snowplow's [loading-process documentation](https://docs.snowplow.io/docs/destinations/warehouses-lakes/loading-process/) for retention and replay options. ## At a Glance - **Read**: Snowplow warehouse loaders (raw enriched events) - **Where the data lives**: Your warehouse or data lake — [BigQuery](https://docs.snowplow.io/docs/destinations/warehouses-lakes/bigquery/), [Snowflake](https://docs.snowplow.io/docs/destinations/warehouses-lakes/snowflake/), [Redshift](https://docs.snowplow.io/docs/destinations/warehouses-lakes/redshift/), [Databricks](https://docs.snowplow.io/docs/destinations/warehouses-lakes/databricks/), [Delta Lake](https://docs.snowplow.io/docs/destinations/warehouses-lakes/delta/), or [Apache Iceberg](https://docs.snowplow.io/docs/destinations/warehouses-lakes/iceberg/) - **Latency**: Daily - **Historical data**: Forward from when the Snowplow pipeline starts loading ## Official Links - [Snowplow destinations overview](https://docs.snowplow.io/docs/fundamentals/destinations/) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # RudderStack Integration How SegmentStream integrates with RudderStack: reads raw event data routed by RudderStack into your data warehouse, layered with identity graph and cross-channel attribution. SegmentStream layers cross-channel attribution and budget optimization on top of RudderStack — reading raw event data routed by RudderStack's warehouse destinations into your data warehouse. ## How the integration works RudderStack is a warehouse-first customer data platform — it collects events from your sources and routes them into the data warehouse of your choice. When RudderStack is configured to write to a warehouse, SegmentStream reads from that same warehouse. [RudderStack warehouse destinations](https://www.rudderstack.com/docs/destinations/warehouse-destinations/) include [BigQuery](https://www.rudderstack.com/docs/destinations/warehouse-destinations/bigquery/), [Snowflake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/snowflake/), [Snowflake Streaming](https://www.rudderstack.com/docs/destinations/warehouse-destinations/snowflake-streaming/), [Redshift](https://www.rudderstack.com/docs/destinations/warehouse-destinations/redshift/), [Databricks Delta Lake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/delta-lake/), [ClickHouse](https://www.rudderstack.com/docs/destinations/warehouse-destinations/clickhouse/), [Postgres](https://www.rudderstack.com/docs/destinations/warehouse-destinations/postgresql/), [SQL Server](https://www.rudderstack.com/docs/destinations/warehouse-destinations/sql-server/), [Azure Synapse](https://www.rudderstack.com/docs/destinations/warehouse-destinations/azure-synapse/), [Materialize](https://www.rudderstack.com/docs/destinations/warehouse-destinations/materialize/), [Amazon S3 Data Lake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/s3-datalake/), [GCS Data Lake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/gcs-datalake/), and [Azure Data Lake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/azure-datalake/). SegmentStream connects to whichever one you use. ## RudderStack is the pipe — SegmentStream is the measurement layer RudderStack routes raw events from your sources to destinations. SegmentStream consumes the same raw events that RudderStack lands in your warehouse, then builds identity graph, cross-channel attribution, and budget optimization on top. No separate event mapping or transformation is required. ## Historical data RudderStack writes warehouse data forward from when the destination is enabled. Sync cadence is configurable up to 24 hours; SegmentStream consumes whichever cadence the destination uses. For replay and retention specifics, see RudderStack's [warehouse destinations documentation](https://www.rudderstack.com/docs/destinations/warehouse-destinations/). ## At a Glance - **Read**: RudderStack warehouse destinations (raw events) - **Where the data lives**: Your warehouse or data lake — [BigQuery](https://www.rudderstack.com/docs/destinations/warehouse-destinations/bigquery/), [Snowflake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/snowflake/) (and [Streaming](https://www.rudderstack.com/docs/destinations/warehouse-destinations/snowflake-streaming/)), [Redshift](https://www.rudderstack.com/docs/destinations/warehouse-destinations/redshift/), [Databricks Delta Lake](https://www.rudderstack.com/docs/destinations/warehouse-destinations/delta-lake/), [ClickHouse](https://www.rudderstack.com/docs/destinations/warehouse-destinations/clickhouse/), [Postgres](https://www.rudderstack.com/docs/destinations/warehouse-destinations/postgresql/), [SQL Server](https://www.rudderstack.com/docs/destinations/warehouse-destinations/sql-server/), [Azure Synapse](https://www.rudderstack.com/docs/destinations/warehouse-destinations/azure-synapse/), [Materialize](https://www.rudderstack.com/docs/destinations/warehouse-destinations/materialize/), or any of [S3](https://www.rudderstack.com/docs/destinations/warehouse-destinations/s3-datalake/) / [GCS](https://www.rudderstack.com/docs/destinations/warehouse-destinations/gcs-datalake/) / [Azure](https://www.rudderstack.com/docs/destinations/warehouse-destinations/azure-datalake/) data lakes - **Latency**: Daily - **Historical data**: Forward from when the warehouse destination is enabled ## Official Links - [RudderStack warehouse destinations](https://www.rudderstack.com/docs/destinations/warehouse-destinations/) ## Related - [All Integrations](https://segmentstream.com/integrations.md) - [Measurement Engine](https://segmentstream.com/measurement-engine.md) --- # Identity Graph: Stitching Users Across Devices and Browsers How SegmentStream resolves fragmented user journeys using deterministic identity stitching — without third-party cookies or statistical inference. --- One person uses three devices. Analytics sees three strangers. Your attribution model scores all three journeys wrong — and every downstream system inherits the error. ## 01 — The Problem ### Your Attribution Is Only as Good as Your Identity Graph Attribution models don't fail because of bad math. They fail because they don't know who they're measuring. ### One person, many cookies A single user's real journey — three touchpoints, three separate cookies, zero connection: 1. **Instagram in-app browser** — user taps an ad. The in-app browser runs in a sandboxed environment with its own cookie. 2. **Safari on iPhone** — later that evening, the user browses your site directly. Safari uses a separate cookie jar from the in-app browser. 3. **Chrome on desktop** — next day at work, the user converts. A third device, a third cookie. Analytics treats each as a separate visitor. Three sessions, three anonymous IDs, zero connection between them. The ad gets no credit. The conversion appears organic. The user's journey is invisible. > **[Interactive: Journey Stitching Toggle]** > Before/after toggle showing the same three visits (Instagram in-app, Safari iPhone, Chrome Desktop) as disconnected raw sessions versus a stitched user journey with linking signals. Toggle switches between "Raw Sessions" view (three separate cards with separate cookie IDs) and "Stitched Journey" view (unified timeline connected by identity signals). ### The in-app browser problem This is the most damaging source of identity fragmentation in paid social. When someone taps an ad on Instagram, Facebook, or TikTok, it opens in the app's embedded browser — a completely separate cookie sandbox from Safari or Chrome. The user sees your landing page, browses a few pages, then switches to their default mobile browser and converts. Analytics records: - **Session 1** (in-app browser): source = `instagram / paid`, no conversion - **Session 2** (Safari): source = `direct / none`, conversion credited Facebook claims the conversion via its own pixel. Your analytics credits "direct." The actual first-touch channel — the ad that introduced the user — gets nothing. This isn't an edge case. On mobile-heavy sites, a significant share of paid social traffic arrives through in-app browsers. > **[Illustration: In-App Browser Attribution Diagram]** > Horizontal user journey showing Instagram opening an in-app browser (cookie_X) which links to Safari (cookie_Y) via shared IP address. Below the journey, two outcome boxes compare "With Identity Graph" (correctly attributes to instagram/paid) versus "Without Identity Graph" (misattributes to direct/none). ### Targeting identity graphs solve the wrong problem The probabilistic identity resolution industry — LiveRamp, UID2, Experian — was built for **ad targeting**. Targeting tolerates false positives. If you show an ad to someone who isn't actually the same person, you wasted a fraction of a cent. The cost of a wrong match is trivial. Measurement cannot tolerate false positives. A single incorrect identity link can: - Credit the wrong channel for a conversion - Inflate one campaign's ROAS while deflating another's - Corrupt budget optimization inputs - Cascade errors through every downstream model Different problems require different tools. Targeting identity graphs optimize for **reach** (link as many identifiers as possible). Measurement identity graphs must optimize for **precision** (only link identifiers when confidence is high). ### Statistical matching error compounds with journey length Identity accuracy degrades exponentially across multi-touch journeys. At 80% identity accuracy (a generous estimate for statistical matching methods): | Touchpoints | Accuracy | |-------------|----------| | 2 | 80% | | 3 | 64% | | 4 | 51% | | 5 | 41% | Each identity link between touchpoints multiplies the error probability. By the time a user has interacted across three devices, statistical identity matching is barely better than a coin flip. This is why SegmentStream uses **only deterministic signals** for identity stitching. Every link between two anonymous IDs requires a shared, verified identifier — a login, an `email_hash`, a `click_id`, or same-network activity within a tight time window. --- ## 02 — The Framework ### Three Resolution Levels. One Deterministic Graph. SegmentStream's identity graph resolves identities at three levels — individual, household, and organizational. Each level uses different signals with different confidence windows. The same signal can create links at different levels depending on context — an IP address links devices for one person, family members at home, or coworkers at the office. All matching is deterministic: shared keys only, no statistical inference. ### Individual identity — one person, multiple devices Four standard signals stitch one person's sessions across devices and browsers: - **User ID** — authenticated login links the current anonymous session to every previous session where the user logged in. Gold standard for identity resolution. 180-day window. - **Email Hash** — email captured at signup, checkout, or newsletter subscription is hashed with SHA-256 and used as an identity key. Same hash on a different device links the sessions. 180-day window. *Status:* supported in the pipeline but not yet deployed in production. Requires adding `?ehash=` parameters to marketing email URLs. - **Click ID** — ad platforms append identifiers (`gclid`, `fbclid`, `ttclid`) to URLs. The SDK captures these on landing and preserves them across internal navigations — including the in-app browser to native browser bridge. 30-day window. - **IP Address** — two anonymous sessions sharing the same IP within the confidence window are linked. Weakest standard signal due to shared networks. 3-day window. Safeguard: IPs shared by more anonymous IDs than the 99th percentile are flagged as outliers and all links via that IP are discarded. ### Email link stitching — the full flow Email hash stitching follows a specific technical flow: 1. **Capture:** User subscribes on Device A. Email is captured and hashed with SHA-256. 2. **Tag:** All marketing email URLs include an `?ehash=abc123` parameter containing the hashed email. 3. **Stitch:** User clicks the email link on Device B. The `ehash` parameter matches the hash from Device A. Both anonymous sessions are now linked to the same universal user. This works because the email hash is the same regardless of which device opens the link. No cookies required. No third-party data. ```html Shop now Shop now ``` > **[Illustration: Email Stitching Flow]** > Horizontal user journey showing Device A (where email is captured and hashed) connecting through an email link (with `?ehash=` parameter) to Device B (where the hash matches). Below, the identity graph links both cookies via the shared email hash. SHA-256 hash only, no raw email stored, 180-day window. ### Household identity — family, shared context Household-level identity links different people within the same family or home — not the same person across devices, but related users who share context. - **IP Address** — family members on the same home Wi-Fi share an IP. Within the 3-day window, their sessions are linked at the household level. - **Click ID** — when someone shares a campaign link with a family member (via WhatsApp, iMessage, or email), the `click_id` propagates with the URL. Both visits carry the same `click_id`, enabling household-level attribution. ### Organizational identity — company, buying committee B2B journeys involve multiple people at the same company: - **Account ID** — multi-tenant B2B products where multiple users share a product account. - **Email Domain Hash** — a junior employee signs up for a trial, a manager requests a demo, a VP visits the pricing page. If all use `@company.com` emails, the domain hash connects them into one organizational journey. Always pass `account_id` alongside `email_domain_hash` so the identity graph can enforce hard boundaries between organizations sharing the same email domain (e.g. agencies using client domains). - **IP Address** — employees at the same office share a corporate IP. Within the 3-day window, their sessions are linked at the organizational level. The 99th-percentile outlier safeguard prevents large offices from creating false merges. > **[Illustration: Signals Grid]** > Visual grid showing identity signals organized by three resolution levels. Individual level: User ID (180d, highest strength), Email Hash (180d, highest), Click ID (30d, high), IP Address (3d, moderate). Household level: Home Network IP (3d, moderate), Shared Link Click (30d, low). Organizational level: Account ID (180d, highest), Email Domain (180d, highest), Corporate Network IP (3d, moderate). Each card shows signal name, key name, time window, strength indicator, and best-for label. ### Custom identity keys Beyond the standard signals, SegmentStream supports arbitrary custom identity keys. Any event property can be designated as an identity key with a configurable confidence window. Examples from production: - `cart_id` — cross-device cart recovery. User starts checkout on one device, completes on another. - `quote_id` — insurance or B2B quotes sent via email, opened on a different device - `checkout_id` — cross-device checkout linking - `loyalty_card_id` — retail loyalty programs - `phone_hash` — SMS marketing attribution - `order_id` — post-purchase support interactions linked to acquisition --- ## 03 — The Algorithm ### Connected Components on Deterministic Edges The identity graph runs as a daily batch pipeline. It takes raw events from GA4 and the SegmentStream SDK, extracts identity keys, finds users that share keys, and groups them into universal user profiles. ### Five-stage pipeline **Stage 1 — Load raw events and extract identity keys** Two data sources feed the pipeline: - **Analytics events** — page views, transactions, and custom events from any analytics platform (GA4, Adobe Analytics, Amplitude, Heap, Segment, or others) - **SegmentStream SDK events** — lightweight client-side pings that capture click IDs and IP addresses across sessions, including the in-app browser to native browser bridge From each event, the pipeline extracts all available identity keys: `user_id`, `email_hash`, `click_id` (`gclid`, `fbclid`, `ttclid`, etc.), `ip_address`, plus any configured custom keys. Each key is paired with the event's `anonymous_id` — the cookie-level identifier. **Stage 2 — Aggregate daily user profiles** All events are grouped by date and `anonymous_id`, collecting the distinct identity keys observed for each visitor on each day. This produces a daily snapshot: "`anonymous_id` X was seen with keys [user_id=123, ip=1.2.3.4, gclid=abc]." **Stage 3 — Calculate user links** The pipeline compares anonymous ID pairs. If two different anonymous IDs share at least one identity key within that key's confidence window, a link is created between them. **Skewed key filtering:** Before creating links, any identity key shared by more than 100 anonymous IDs is discarded entirely. All links through that key are dropped. This prevents a single corporate IP address or shared WiFi network from merging hundreds of unrelated users. **Stage 4 — Calculate connected components** The pipeline takes all pairwise links and finds connected components — groups of anonymous IDs that are transitively linked. If identities are transitively connected, they are merged into a single component: - A is linked to B (via shared `ip_address`) - B is linked to C (via shared `email_hash`) - Therefore A, B, and C are all the same user — even though A and C share no direct identity key **Component size cap:** Connected components larger than the 99th percentile are flagged as outliers and discarded. Components that large indicate shared infrastructure or a data quality issue, not real users. **Hard constraints:** If two anonymous IDs have different `account_id` values, they are never linked, regardless of other shared keys. This prevents cross-account contamination in multi-tenant B2B products. **Stage 5 — Output to users table** Each anonymous ID is mapped to a `universal_id` — the identifier of the canonical user. The universal ID is the anonymous ID with the earliest `first_visit` timestamp in the component. This ensures the user's identity anchors to their oldest known session. The output is written to the `users` table in the customer's BigQuery warehouse. Every downstream system — attribution, scoring, reporting — joins against this table to resolve anonymous IDs to universal users. > **[Illustration: Identity Pipeline Flow]** > Animated 3-part horizontal flow. Left: website wireframe generating cookie events. Center: identity graph circle where cookies connect and merge via shared identity parameters (ip_address, email_hash, user_id, click_id — one per round). Right: resolved users table accumulating universal_id rows. Four rounds demonstrate each stitching parameter in sequence. ### Key properties **Deterministic only.** Every link requires a shared, verified identity key. No statistical inference, no behavioral similarity, no device fingerprinting. **Transitive resolution.** If A links to B and B links to C, all three get the same universal ID — even if A and C share no direct key. This is what makes connected components powerful: identity signals compound across touchpoints. **Warehouse-native.** The pipeline reads from and writes to the customer's BigQuery project. No data leaves their warehouse. No external identity graph services. The customer owns their identity data. **Batch processing.** The pipeline runs daily. Approximately one day of lag between an event occurring and it being reflected in the identity graph. This is a deliberate tradeoff — batch processing enables the full connected-components algorithm at scale. ### Privacy by design **Non-consent users are invisible.** Users who decline consent receive a `non-consent-{uuid}` anonymous ID that changes on every page load. These IDs are never linked to anything — the identity graph literally cannot see them. **No raw emails stored.** Only SHA-256 hashes of email addresses are used as identity keys. The raw email is never written to the identity graph pipeline. **All data stays in customer's BigQuery.** The pipeline reads from and writes to the customer's BigQuery project — no raw data or PII is transferred outside their infrastructure. --- ## 04 — In Practice ### See It Work #### Querying the identity graph via MCP Two MCP tools expose the identity graph directly. `get_identity_graph_statistics` returns signal coverage, stitching rates, and device counts for a given project. `get_user_journey` resolves a specific user's cross-device timeline — showing every anonymous session, the signals that linked them, and the full attribution path. > **[Interactive: Identity Graph MCP Terminal]** > Claude Code-style terminal with two rotating scripts. Script 1 runs `get_identity_graph_statistics` for Purchase conversion: shows a signal coverage table (click_id 56.2%, ip_address 100%, email_hash 73.2%, user_id 88.4%), converted user count (1,228), cross-device rate (47.6%), and average 4.6 devices per stitched user. Script 2 runs `get_user_journey` for the most recent conversion: shows a 3-device journey (iPhone in-app via instagram/paid_social, iPhone Safari via IP match, Desktop Chrome via user_id match) ending in a $248 purchase, with attribution impact analysis comparing "without" (direct/none gets 100% credit) versus "with" (instagram/paid_social credited as first touch). #### Production benchmarks Median stitching rates across production projects, grouped by the number of deterministic identity parameters passed: > **[Interactive: Stitching Stats]** > Animated bar chart showing two benchmarks. "2 parameters" (click_id + IP): 35% median stitching rate. "3 parameters" (click_id + IP + user_id/email_hash): 67% median stitching rate. Bars animate on scroll with count-up numbers. #### What the numbers tell you **Adding User ID lifts stitching from 35% to 67%.** Projects passing only `click_id` and `ip_address` achieve a median 35% stitching rate across converted users. Adding a third deterministic signal — typically `user_id` or `email_hash` from a backend integration — pushes that to 67%. A +32 percentage point improvement from a single additional parameter. **"2+ devices" means 2+ anonymous sessions stitched.** This includes true cross-device journeys (phone to laptop), but also same-device fragmentation: in-app browsers, cleared cookies, incognito sessions. All of these create orphaned sessions that need stitching regardless of cause. **Coverage quality matters as much as key count.** Projects with GDPR consent constraints see `ip_address` and `click_id` coverage drop to ~70%, reducing their effective stitching rate even with 3 parameters. The benchmark assumes well-integrated projects with standard cookie consent. #### The retargeting test Here's a practical way to validate your identity graph is working: Check first-touch attribution on retargeting and email campaigns. In a well-functioning identity graph, these should show **close to zero first-touch conversions**. Why? Retargeting and email reach people who already visited your site. If the identity graph correctly stitches the retargeted user's new session to their original visit, first-touch credit goes to the channel that originally introduced them — not to "retargeting" or "email." If retargeting shows significant first-touch conversions, it means the identity graph failed to link the retargeted session to the user's original visit. They look like new users when they're not. Ideal result: zero first-click conversions for retargeting campaigns. That means identity stitching is doing its job — these campaigns are being correctly recognized as re-engagement, not acquisition. > **[Interactive: Attribution Split View]** > Three-column comparison table showing 6 channels (Paid Social, Paid Search, Organic, Retargeting, Email, Direct) with first-touch conversion counts "Without Identity Graph" versus "With Identity Graph." Key shifts: Retargeting drops from 18 to 5, Email drops from 12 to 2, Direct drops from 51 to 38. Paid Social rises from 26 to 41, Paid Search from 22 to 32, Organic from 18 to 29. Numbers animate with count-up on scroll. --- ## 05 — Common Mistakes ### What Goes Wrong and How to Avoid It #### Using statistical matching for measurement decisions Statistical matching — behavioral fingerprinting, device graph lookups, inferred identity links — is designed for ad targeting where false positives are cheap. When applied to measurement, false links corrupt attribution data silently. You won't see an error. You'll see wrong numbers that look plausible. **Fix:** Use only deterministic identity signals for measurement. Accept a lower stitching rate in exchange for trustworthy data. #### Not installing the SegmentStream SDK Without the SDK, `click_id` preservation doesn't work. The SDK captures `gclid`, `fbclid`, `ttclid`, and other click parameters on landing and preserves them across internal page navigations. Without it, these values are lost after the landing page. **Fix:** Install the SDK on all pages. It's a single JavaScript snippet. #### Not sending User ID on every authenticated page load A common implementation mistake: sending `user_id` to GA4 only on the login page. If a user is already logged in and visits other pages, those sessions don't carry `user_id` — and the identity graph can't link them. **Fix:** Set `user_id` in the GA4 config on every page load for authenticated users, not just the login event. #### Inconsistent email normalization before hashing If one system hashes `User@Example.com` and another hashes `user@example.com`, the SHA-256 outputs will differ. The identity graph sees two different keys and can't link them. **Fix:** Normalize before hashing — lowercase, trim whitespace, strip dots from Gmail addresses (if applicable). Apply the same normalization everywhere email hashes are generated. #### Expecting real-time stitching The identity graph runs as a daily batch pipeline. If a user visits on Device A in the morning and Device B in the afternoon, the stitching won't appear until the next day's pipeline run. **Fix:** Design your workflows around daily data freshness. Identity stitching is not suited for real-time personalization — it's built for accurate measurement and reporting. #### Ignoring consent compliance Non-consent users receive a rotating anonymous ID (`non-consent-{uuid}`) that changes every page load. This is by design — the identity graph cannot track users who haven't consented. If consent rates are low, your stitching rate will be structurally limited. **Fix:** Optimize your consent flow. Higher consent rates directly improve identity graph coverage. There is no technical workaround — and there shouldn't be. #### Skipping the corporate IP safeguard If you're seeing unusually high stitching rates (40%+), check for corporate IP merges. A shared office IP can link hundreds of unrelated users. The pipeline's default safeguard — discarding identity keys shared beyond the 99th percentile of anonymous IDs — catches most cases, but misconfigured custom keys can still cause issues. **Fix:** Monitor the key combination breakdown in `get_identity_graph_statistics`. If a single IP or custom key is driving a disproportionate share of stitching, investigate. #### Agency cross-account contamination in B2B The `account_id` signal is primarily used by B2B and SaaS products for organizational identity — linking multiple employees within the same company into a single buying committee journey. The identity graph treats `account_id` as a hard constraint: two anonymous IDs with different `account_id` values are never linked, regardless of other shared signals. **Example:** john@accenture.com visits your site under `account_id` "accenture-client-a" and michael@accenture.com visits under `account_id` "accenture-client-b". Their `email_domain_hash` is identical (both @accenture.com), but because their `account_id` values differ, the identity graph will not stitch them. The hard constraint overrides the shared domain hash — cross-account contamination is blocked. This creates a specific risk for agencies and consultancies (Dentsu, GroupM, Accenture) whose employees access multiple client accounts. If the CRM assigns a shared or agency-level `account_id` rather than a client-specific one, the identity graph will merge sessions across unrelated client accounts. The hard constraint only works when the boundary is correct upstream. **Fix:** Ensure each `account_id` maps to a single end-client account, not an agency umbrella. Always pass `account_id` alongside `email_domain_hash` — this lets the identity graph automatically enforce hard boundaries between organizations, even when employees share the same email domain. Monitor component sizes via `get_identity_graph_statistics` — unusually large components often indicate cross-client leakage from shared account identifiers. --- # Cross-Channel Attribution First-click attribution and predictive conversion maturation — two methods that cross-validate to produce click-time revenue you can optimize against. --- Every attribution model has a bias. Most of them bias in the same direction — toward the channels that touched the user last, not the ones that introduced them. That bias isn't a bug. It's the business model. ## 01 — The Direction Problem Attribution models disagree on how to distribute credit. But they almost all agree on direction: credit flows toward the bottom of the funnel. The channel that touched the user last gets the most. The channel that introduced them gets the least — or nothing. This isn't a quirk of one model. It's a pattern across three distinct mechanisms, each independently tilting credit the same way. - **Retargeting Machine** — retargeting ads claim credit for conversions that were already happening. - **Brand Search Cannibalization** — brand ads intercept organic traffic already navigating to your site. - **Identity Fragmentation** — broken identity makes returning visitors look like new ones. ### The retargeting machine It starts the moment someone visits your site. A platform pixel fires. The user is added to retargeting audiences across every ad network with a pixel on your page — Google Display, Meta, TikTok, your DSP — before your exclusion audiences even update. From here, the credit theft escalates in two stages: **Click retargeting.** The user sees a retargeting ad, clicks it, and converts. The retargeting platform claims the conversion. The prospecting ad that introduced this person — the only channel that did something no other channel could have done — gets no credit. > Organic visit triggers pixel. User is retargeted within minutes. Retargeting claims the conversion. The discovery channel gets nothing. **Impression retargeting (post-view).** The user doesn't even click. They see an ad — or more likely, they scroll past one — and convert later on their own. The platform claims a "view-through" conversion. The user never engaged. The platform just happened to show an ad to someone who was already going to buy. Post-view is worse than click retargeting because the bar for claiming credit is zero interaction. A large share of display ad impressions are never actually seen — the user never scrolled to them, or the ad loaded below the fold. Yet platforms count post-view conversions for all of them. Platform-reported conversions routinely sum to several times the actual number. This is not a bug. It's how post-view attribution is designed to work. > One organic conversion, three platform claims. Post-view attribution requires zero interaction — every platform with a pixel claims credit for the same purchase. The result: reported conversions routinely sum to several times the actual number. The attribution window is the tell. Facebook shortened its default post-view window from 28 days to 1 day — not because 1 day is more accurate, but because a 7-day window would attribute nearly everything to social, and that doesn't look realistic. The window was shortened for optics, not accuracy. A calibrated lie. Here is the question this raises: if you have platform pixels on your website, every returning visitor gets retargeted. What conversion *isn't* post-view retargeting? Standard attribution often shows retargeting ROAS of 10x or higher. Incrementality testing consistently reveals the true lift is a fraction of that — sometimes negative. The gap exists because retargeting excels at claiming credit for conversions that were already happening. > Last-click concentrates credit in brand search and retargeting. Post-view smears credit across everything — channels with minimal spend show substantial attribution because everyone scrolls past ads. ### Brand search cannibalization When someone searches your brand name, they already know you exist. They are navigating to your site, not discovering it. The organic result is right below the ad — same destination, one scroll away. Without the ad, the click goes to the organic link. The visit still happens. The conversion still happens. Brand search ads intercept this intent and claim credit for it. Attribution records the ad click as the converting touchpoint. But the ad didn't create the visit — it intercepted a visit that was already happening. The reported ROAS is high because the denominator is small (brand keywords are cheap) and the numerator is large (these users were already going to buy). The actual incremental impact — conversions that would not have happened without the ad — is a fraction of what's reported. Geo-testing confirms this reliably. Pause brand ads in a set of regions, keep them running everywhere else, and compare. The pattern is always the same: nearly all of the paid brand traffic shifts to organic. The conversions don't disappear. They just arrive through a different link. The gap between reported brand search performance and actual incremental impact is consistently enormous — often an order of magnitude. > Paid brand ads intercept clicks that were already heading to the organic result. Same destination — the ad didn't create the visit, it intercepted one. ### Identity fragmentation Any time the converting session can't be linked back to the session that introduced the user, credit goes to the wrong channel. The gap between the ad click and the conversion creates opportunities for misattribution. - **In-app browsers** — Instagram, TikTok, and other apps open links in embedded browsers with separate cookie jars. - **Cross-device** — user clicks an ad on mobile, converts on desktop. - **Cross-browser** — user browses in Safari at work, converts in Chrome at home. - **ITP cookie expiry** — Safari caps first-party cookies at 7 days; returning after a week looks like a new visitor. - **Incognito / private browsing** — session starts with zero history. > One user, three cookie jars. The ad click that introduced the user gets zero credit because analytics sees three separate visitors. This problem is covered in depth in the [Identity Graph](/measurement-engine/identity-graph.md) page. > Three independent phenomena. Each one systematically over-credits the channels that touched the user last, not the ones that introduced them. This isn't noise. It's a tilt. --- ## 02 — Five Properties of Trustworthy Attribution Before choosing an attribution model, define what you expect from one. Not every model needs every property. But knowing which properties a model has — and which it lacks — prevents the common mistake of trusting a number because it came from something called "data-driven." The table below scores seven attribution models against five properties. Each property is a standalone requirement — a question you should be able to answer about any model before trusting its output. **Traceable** — A conversion must be traceable to a specific session: one click, one timestamp, one landing page. Without traceability, you cannot investigate why a specific conversion was attributed the way it was. Traceability is the foundation of debugging and auditing. **Explainable** — The attribution logic must be explainable in one sentence to a non-technical stakeholder. If you cannot explain it simply, you do not understand it well enough. Complexity that cannot be communicated cannot be challenged. **Directionally Neutral** — Does the methodology systematically over-credit one end of the funnel? A directionally neutral methodology has no structural tilt — its errors are either random or self-correcting, rather than consistently favoring acquisition or re-engagement channels. **Auditable** — The raw data and attribution logic must be inspectable in your own data warehouse. If you only see the output but not the inputs or the calculation, you are trusting a black box. Auditability means you can reproduce the result independently. **Deduplicated** — Each conversion must be counted exactly once across all channels. If multiple systems independently claim credit for the same conversion, the total attributed conversions will exceed actual conversions. Deduplication requires a single source of truth. > **[Interactive: Attribution Comparison Table]** > HTML table scoring seven attribution models (Ad Platform Post-Click, Ad Platform Post-View, Ad Platform "Incremental", Last-Click, First-Click, Algorithmic MTA, "Cookieless") against five properties (Traceable, Explainable, Dir. Neutral, Auditable, Dedup.). Each cell shows a pass (filled circle), partial (half circle), or fail (X circle) icon. First-Click is the only model that passes all five properties and is highlighted with an accent row. > **[Interactive: Attribution Rationale]** > Seven collapsible accordion cards, one per attribution model. Each expands to show the scoring rationale for all five properties with pass/partial/fail icons. Explains why each model earned its score. ### Same journey, three models Consider a real conversion path: Meta Prospecting ad, then an organic return visit, then Google retargeting, then TikTok retargeting, then a brand search click, then conversion. > **[Interactive: Journey Three Models]** > A five-touchpoint journey (Meta Prospecting, Organic Return, Google Retargeting, TikTok Retargeting, Brand Search) shown under three attribution models via a segmented control. Last-Click gives 100% to Brand Search. Algorithmic MTA distributes credit across all touchpoints (18%, 12%, 22%, 25%, 23%) — the discovery channel gets less than each retargeting channel. First-Click gives 100% to Meta Prospecting. Horizontal credit bars animate when switching models. Each model includes a verdict explaining its bias. ### The triangulation fallacy A common response to broken attribution: "Use multiple models and triangulate." The logic sounds reasonable. GPS uses multiple satellites. Surely multiple attribution models converge on truth. The analogy fails. GPS satellites measure the same physical reality — the speed of light is constant, the satellite positions are known, and the signal propagation is governed by physics. Attribution models don't measure reality. They impose different allocation rules on the same incomplete data. Averaging a compass that points north and a compass that points south doesn't give you east. It gives you the illusion of a direction. --- ## 03 — Why First-Click Is Closest to Incrementality The question attribution should answer isn't "which channel touched the user last?" or "how should credit be distributed fairly?" It's: "which channel introduced this person for the first time?" That question — who introduced the user — is the closest any click-based attribution method can get to incrementality. A channel that brings someone to your site for the first time has done something no other channel in the path could have done. Every subsequent touchpoint — retargeting, email nurture, brand search — requires that first visit to exist. ### The logic chain 1. **Incrementality** measures whether a conversion would have happened without marketing intervention 2. The most incremental touchpoint is the one that **introduced the user** — without it, no subsequent touchpoints exist 3. The channel that introduced the user is the **first click** First-click attribution is not perfect incrementality. But it is the only rule-based attribution method whose bias — crediting discovery — aligns with the question marketers actually need answered: "Where are my new customers coming from?" ### Identity graph prerequisite First-click attribution is only as good as your ability to identify who the user is across sessions. If the identity graph fails to stitch a retargeted session back to the user's original visit, first-click credit goes to the retargeting channel — the same error every other model makes. This is why [the identity graph](/measurement-engine/identity-graph.md) is a prerequisite, not an optional add-on. Sessions are grouped by `universal_id` (post-identity-graph), not by `anonymous_id`. The first click is the first session for the resolved user, not the first session for a cookie. ### Attribution window First-click does not mean first-ever. The attribution window is configurable: 30 days by default, up to 90 days maximum. If a user's first visit was 120 days ago, the attribution window has closed. Their next visit starts a new attribution cycle. This prevents the absurdity of crediting a Facebook ad from six months ago for today's conversion. The window defines how far back "first" can mean. Attribution accuracy matters because every downstream decision — budget allocation, marginal ROAS, channel sequencing — compounds from it. Wrong attribution at the source produces wrong optimization everywhere else. This is the attribution methodology the SegmentStream Measurement Engine implements: first-click credit assignment on identity-resolved journeys, with click-time reporting and maturation prediction. The sections below describe how. --- ## 04 — Click-Time vs Conversion-Time Reporting Attribution determines which channel gets credit. But there is a separate question that distorts the numbers just as badly: *when* does that credit land in your report? Most analytics platforms default to conversion-time reporting — revenue appears on the date the conversion happened. This creates a problem whenever spend and conversions don't land in the same period. Cut your budget this month and conversions from last month's high spend inflate this month's ROAS. Increase spend and conversions haven't caught up yet — ROAS craters. Shift budget between channels and the lag makes the old channel look better and the new one look worse. Any budget fluctuation creates the distortion. ### The seasonal illusion The distortion is easiest to see when spend varies across months. A business runs campaigns from April through September, scaling up for peak season and down after. True ROAS is 3.0 across every month — same campaigns, same performance. But their analytics shows something different: > **[Interactive: Seasonal Toggle]** > Toggle between Click-Time and Conversion-Time views of the same campaign data (April–September). Click-Time table shows consistent ROAS 3.0 across all months. Conversion-Time table shows distorted ROAS: April 0.6, May 1.2, June 4.5, July 4.0, August 2.2, September 10.2 — with "illusion" labels explaining the wrong conclusion each number would trigger. Same campaigns, same performance, opposite conclusions. Conversion-time reporting attributes revenue to the month the conversion happened, not the month the click happened. Clicks in April drive conversions in June. The result: April looks like a failure (ROAS 0.6) and the team concludes "campaigns underperforming — cut budget." September shows ROAS 10.2 and the team concludes "best month ever! Keep going!" Same campaigns. Same performance. Opposite conclusions — both wrong. Click-time reporting with maturation prediction fixes this — accurate ROAS within days, not months. ### Click-time reporting Click-time reporting fixes this by stamping revenue to the date of the original click, not the date of the conversion. Every conversion produces two records: one at the conversion date (for finance), one at the click date (for marketing). Click-time reporting answers: "What is the ROAS of the clicks I paid for in April?" — regardless of when those clicks converted. Conversion-time reporting answers a different question: "How much revenue was recognized in April?" Both are valid. But only click-time ROAS tells you whether your April campaigns worked. ### The maturation tradeoff Click-time reporting has a catch: recent periods are always incomplete. Clicks from two weeks ago are still converting. ROAS for recent cohorts is artificially low because conversions are still arriving. This is where conversion maturation prediction comes in. The system projects what the final ROAS will be based on the maturation pattern of older, fully-converted cohorts. The result: you can read click-time ROAS for last week without waiting a month for conversions to finish arriving. > The reporting window is invisible infrastructure. Get it wrong and every ROAS number, every budget decision, every scaling call is built on a mirage. Click-time reporting with maturation prediction is the fix. --- ## 05 — The Infrastructure First-click attribution is a rule. Implementing it correctly requires infrastructure that most analytics stacks don't provide. The SegmentStream Measurement Engine runs three capabilities underneath the attribution output, each solving a specific measurement problem. ### Conversion maturation prediction Click-time reporting has a lag problem. Clicks from this week have not finished converting yet. ROAS for recent periods is artificially low because conversions are still arriving. [Predictive Attribution](/measurement-engine/predictive-attribution.md) fills this gap by projecting the final result of immature click cohorts. In cross-channel attribution, the important point is not the model internals; it is that recent click-time ROAS must include a maturation adjustment before it is used for budget decisions. > **[Illustration: Maturation Curve]** > Horizontal bars showing weekly conversion maturation over 6 weeks: Week 1 (38%), Week 2 (65%), Week 3 (82%), Week 4 (91%), Week 5 (96%), Week 6 (99%). Each bar has a filled portion (observed) and empty portion (remaining expected). A vertical dashed "You are here" line at Week 3 marks 82% tracked. Earlier weeks are fully matured; later weeks show observed counts understating true performance. ### Noise smoothing The smaller the campaign, the noisier the signal. A campaign averaging 2 conversions per day — 14 per week — swings 200% from pure randomness. That isn't a performance change. It's a coin flip. > **[Illustration: Noise Sample Size]** > Three horizontal range bars centered on true CPA $60, showing variance at different conversion volumes. 50 conv/day: $54–$66 (±10%), tight range highlighted in accent. 5 conv/day: $30–$90 (±50%), medium range. 1 conv/day: $0–$180 (±100%), full-width range. Bars expand outward from the true CPA center line on scroll. For most campaigns, neither daily nor weekly data has enough conversions to be statistically meaningful. "CPA spiked 150% yesterday, pause the campaign." "ROAS dropped 40% this week, cut budget." Both are reactions to noise, not signal. The campaign didn't change. The sample is just too small for either reporting window. Noise smoothing applies sliding-window normalization with anomaly detection. Spikes above 1.5x the window median and drops below 0.5x are flagged and dampened. The maturation confidence score penalizes high-variance cohorts — if the data is noisy, the prediction says so. > **[Illustration: Noise Smoothing]** > Two-line overlay chart. Raw daily CPA (gray, jagged line) swings wildly across 14 data points. 7-day rolling average (accent/indigo, stable line) barely moves. The visual gap between the lines IS the noise. End labels show yesterday's raw CPA vs. the 7-day average. The campaign didn't change — the sample size is too small for daily reporting. ### Consent modeling Users who decline cookie consent are invisible to attribution. The identity graph cannot see them — by design. But they still click ads and they still convert. Ignoring them systematically underreports every channel's true performance. The Measurement Engine fills this gap using bucket-level statistical redistribution. For non-consented clicks, the system captures: date/time, geolocation, user agent, landing page product, and traffic source. For non-consented conversions: date/time, geolocation, user agent, and purchased product. These signals are matched at the bucket level — channel by geo by product category by time window — to redistribute unattributed conversions proportionally across channels. > **[Interactive: Consent Modeling Table]** > HTML table showing four sources (Google Ads, Meta Ads, Direct, Unattributed) with columns for Clicks, Sessions, Conversions, Consent %, and Modeled conversions. Below the table, an animated SVG shows the bucket matching process: left-side click dimensions (Date & Time, Geolocation, User Agent, Landing Page Product) and right-side conversion dimensions (Date & Time, Geolocation, User Agent, Purchased Product) converge through merge brackets into a central "Bucket Match" box. The limitation is explicit: consent-modeled conversions have no individual customer journey. They are statistical estimates, not deterministic links. The system labels them accordingly so downstream reports can distinguish between tracked and modeled conversions. --- ## 06 — Calibrated Imperfection ### Two opposing forces First-click attribution has a built-in balancing mechanism that most people miss. Two forces pull credit in opposite directions: **Force 1: Stitched journeys push credit up.** When the identity graph successfully links a returning visitor to their original session, first-click credit flows to the channel that introduced them — by definition an acquisition or awareness channel. Better stitching means more upper-funnel credit. **Force 2: Unstitched journeys push credit down.** When the identity graph fails to link sessions, a returning user's retargeting or brand search visit looks like a first visit. Credit flows to a lower-funnel channel. Worse stitching means more lower-funnel credit. These forces pull in opposite directions. The stitching rate is the dial. At real-world stitching rates, the result is approximate directional neutrality — not because the system is tuned for it, but because the two error modes naturally counterbalance. ### The retargeting test You can measure the balance directly. Check the first-click conversion rate on retargeting campaigns. Retargeting reaches people who already visited your site. If the identity graph correctly stitches the retargeted session to the original visit, first-click credit goes to the channel that originally introduced the user — not to retargeting. Near-zero retargeting credit means Force 1 dominates: stitching is working. If retargeting shows significant first-click conversions, Force 2 is active: the identity graph is failing to link sessions. The retargeting first-click rate tells you where you sit on the spectrum between upper-funnel bias and lower-funnel leakage. > **[Interactive: Attribution Split View]** > Three-column comparison table showing 6 channels (Paid Social, Paid Search, Organic, Retargeting, Email, Direct) with first-touch conversion counts "Without Identity Graph" versus "With Identity Graph." Key shifts: Retargeting drops from 18 to 5, Email drops from 12 to 2, Direct drops from 51 to 38. Paid Social rises from 26 to 41, Paid Search from 22 to 32, Organic from 18 to 29. Numbers animate with count-up on scroll. ### The 100% stitching counter A natural objection: "So first-click only works because stitching is imperfect? At 100% stitching, it would just be an upper-funnel bias machine." The opposite is true. At 100% stitching, every journey is fully resolved. Every retargeted session is linked to its original visit. Every brand search click is traced back to the campaign that introduced the user. In that scenario, first-click credit goes to the channel that genuinely introduced the person — pure incrementality credit. What looks like "upper-funnel bias" at perfect stitching is actually the correct answer: the acquisition channel did introduce the user, and every subsequent touchpoint was re-engagement. The directional neutrality at current stitching rates is a fortunate practical property. The theoretical endpoint is even better. > Perfect attribution doesn't exist. The choice is between a model with a known, disclosed, testable upper-funnel bias — and models with hidden, undisclosed, untestable biases that systematically favor the channels selling the most ads. --- ## 07 — See It Work Everything described above runs inside the SegmentStream MCP server. Two queries show cross-channel attribution in practice. The first pulls a channel-level attribution report with maturation predictions. The second traces a single user's journey to show which touchpoint received first-click credit and why. > **[Interactive: Attribution Terminal]** > Claude Code-style terminal with two tabs. Tab 1 "Attribution Report": prompt "Show first-click attribution by channel for last 30 days with maturation" calls `get_attribution_model` and `get_report_table`. Shows a 6-channel table (Meta Prospecting, Google Search, TikTok Ads, Google Retargeting, Meta Retargeting, Brand Search) with columns for Costs, observed Revenue, observed ROAS, projected Revenue, and projected ROAS. Insight: Meta Prospecting is 1.3x observed but 3.2x projected; TikTok looks unprofitable at 0.9x but projects to 2.4x; retargeting stays below 1.0x projected. Action: shift $6K from retargeting to Meta Prospecting. Tab 2 "User Journey": prompt "Show the attribution path for the most recent conversion" calls `get_user_journey`. Shows a 3-device, 4-session journey (iPhone Instagram in-app via meta/paid_social with fbclid, iPhone Safari via IP match, Desktop Chrome via user_id match with google/retargeting and google/brand_search) ending in a $312 purchase. Attribution result: First-click, 100% credit to meta/paid_social. ### What the report shows The attribution report combines three layers: first-click credit assignment (which channel introduced the user), click-time reporting (revenue attributed to the click date, not the conversion date), and maturation prediction (projected final ROAS for immature cohorts). The user journey query exposes the full resolution path: every anonymous session, the identity signals that linked them into a single `universal_id`, the first session that received attribution credit, and the traffic source of that session. ### Validating the output Every number in the attribution report is auditable. You can trace a conversion to a specific `universal_id`, see every session in that user's journey, verify which session was first within the attribution window, and confirm the traffic source recorded for that session. The data lives in your BigQuery warehouse. You can query it directly. The attribution logic is deterministic: given the same inputs, it produces the same outputs. There is no black box. --- # Predictive Cross-Channel Attribution > Click-time attribution is correct but incomplete — you're making decisions on data that won't be final for weeks. Predictive Attribution closes that window. Click-time attribution is correct but incomplete — you're making decisions on data that won't be final for weeks. Predictive Attribution closes that window. *10 min read · April 2026* ## 01 — Click-Time Attribution Is Correct but Incomplete The [Cross-Channel Attribution](/measurement-engine/cross-channel-attribution.md) page explains why click-time reporting is the right foundation: revenue is attributed to the date of the click, not the date of the conversion. This eliminates the seasonal distortions that make conversion-time ROAS unreliable. But click-time reporting has a structural gap. Conversions arrive with a delay. A click today might not convert for days or weeks. Until that conversion arrives, the click looks like it produced nothing. Recent cohorts always look understated because their conversions are still in transit. ### The delay is not small For a typical retailer, only 40% of conversions happen within 7 days of the click. Another 21% arrive between days 8 and 30. The remaining 39% take 30 days or longer. When you look at last week's campaigns, you are seeing less than half the picture. This creates an impossible tradeoff. You can wait until the data matures — but by then the optimization window has closed. Campaigns have been paused, budgets have shifted, and the feedback is stale. Or you can act on incomplete data — but every decision is based on numbers that will look different in three weeks. Neither option works. Waiting loses you optimization time. Acting early gives you the wrong answer. The only way out is prediction: project what the final numbers will look like based on the maturation pattern of older, fully-converted cohorts. > Click-time attribution is correct but incomplete. You are making decisions on data that will not be final for weeks. Predictive Attribution closes that window. --- ## 02 — How Prediction Works Predictive Attribution works at the individual user level. For every visitor who has not yet converted, a machine learning model estimates the probability they will convert within the maturation window. These probabilities are summed across all unconverted users to produce a projected conversion total for each campaign, channel, and time period. ### Maturation window The system automatically calculates how long conversions take. It analyzes historical data to find the maturation window — the number of days it takes for 95% of conversions to be reported. This window varies by conversion type: early-funnel events like demo bookings might mature in 7 days, while purchases might take 60 or 90 days. The maturation window determines how far the model looks forward. For a conversion that matures in 30 days, the model projects for the full 30 days. For the most recent 4–5 weeks of data, this projection fills the gap between what has been observed and what will ultimately arrive. > **[Illustration: Predictive Maturation]** > Maturation curve showing observed conversions vs projected final total for immature cohorts. Solid line for observed data, dashed line for projected portion. ### User-level scoring The model scores each unconverted user based on their engagement signals: page views, return visits, active days, device type, geographic region, traffic source, and campaign characteristics. Users who visit once and never return quickly fade to near-zero probability. Users who come back, browse more pages, and engage deeper carry a higher projected conversion probability. This user-level approach is what makes the projection sensitive to traffic quality. If a campaign shifts to higher-intent audiences, the model sees the behavioral difference in real time — it does not rely on historical averages. ### Calibration Raw model scores are calibrated so that the sum of projected conversions matches expected totals. The calibrator is trained on recent mature data — periods old enough that actual conversions are known. This ensures the projection is not just directionally correct but numerically accurate. > The model scores every unconverted user individually. The projection is not a statistical average — it reflects the actual behavior of the people on your site right now. --- ## 03 — Adaptation Over Time Conversion patterns are not static. The time between a click and a conversion shifts with seasons, promotions, and changes in audience mix. The projection must adapt — without manual intervention. ### The model retrains on recent data The model retrains monthly on the most recent mature data — periods where conversions have fully arrived. This means the model always reflects current conversion behavior. If users start converting faster during a sale, the next training cycle captures that shift. If a new campaign brings a different audience profile, the model learns from their behavior. ### Campaign-level sensitivity Because the model scores individual users based on their engagement, it naturally adapts to changes in traffic quality at the campaign level. A prospecting campaign that brings new, unfamiliar visitors will show lower projected conversion probabilities than a retargeting campaign bringing back engaged users. This distinction happens automatically — no manual campaign classification required. > **[Illustration: Seasonal Maturation]** > Seasonal maturation shift. Peak season compresses the curve, pre-season stretches it. ### Why this matters for decisions Without adaptation, projections trained on Q3 data would systematically mispredict Q4 campaigns. The model's regular retraining on recent data ensures that seasonal and audience shifts are absorbed, not ignored. --- ## 04 — What You See in Reports Predictive Attribution adds projected metrics alongside observed ones in every attribution report. You always see both — what has been tracked and what the model expects. ### Conversions (Incl. Projected) The primary metric. It shows observed conversions plus projected conversions that have not happened yet but are statistically likely to occur. For example, if a campaign shows 10 observed conversions and the model projects 3 more, the metric shows 13. For fully matured cohorts — clicks old enough that conversions have stabilized — the projected portion is zero. The metric equals observed conversions. For recent cohorts, the projected portion fills the gap. The further back you look, the less projection matters. ### Derived metrics CPA (Incl. Projected) and Conversion Rate (Incl. Projected) are calculated from the projected conversion total. This gives you cost efficiency and conversion rate that reflect what the campaign will deliver once conversions finish arriving — not just what has been recorded so far. Projected value metrics extend the same logic to revenue. Conversion Value (Incl. Projected) combines observed conversion value with the expected value of conversions that have not arrived yet. ROAS (Incl. Projected) is then calculated from projected conversion value divided by cost, so recent click cohorts can be evaluated on expected final revenue, not only the revenue recorded so far. ### How projection fades over time The projection is largest for the most recent period and shrinks as conversions arrive. A campaign from last week might show 60% observed and 40% projected. The same campaign two weeks later shows 85% observed and 15% projected. By the time the maturation window closes, the projection is gone — only hard data remains. > The projection fills the gap in recent data. As conversions arrive, projection is replaced by reality. The total stays stable. --- ## 05 — Validation Predictions are only useful if they are accurate. Before enabling projections for any project, the system runs backtests — and you can rerun them at any time. ### How backtesting works Pick a period far enough in the past that the maturation window has fully closed. Look at the data as it appeared at the end of that period — say, 70 observed conversions and a projection of 100. Then look at today's actual total for the same period. If the mature number is close to 100, the projection was accurate. This validation runs at multiple levels: overall, by ad platform, and by campaign. The projection is not just correct on average — it is validated where you make decisions. > **[Illustration: Predictive Validation]** > Backtesting accuracy. Predicted vs actual ROAS after full maturation. ### Projection accuracy by campaign type Not all campaigns project equally well. Retargeting campaigns convert quickly — most conversions arrive within days, leaving little to project. Prospecting campaigns convert slowly — the model fills a larger gap, and accuracy depends on how stable the conversion pattern is. The backtesting makes this transparent. You can see which campaigns benefit most from projection and which already have near-complete data without it. > You do not have to trust the model. You can measure it — and every mature cohort adds to the track record. --- ## 06 — See It Work Predictive Attribution runs inside the SegmentStream MCP server. Ask for projected conversions or projected ROAS in natural language and the server returns the full picture: observed conversions, projected conversions, observed value, projected value, and projected efficiency metrics for every channel — including cohorts that are still maturing. ### What the report shows The MCP server adds projection to the standard cross-channel output. For each channel and campaign: observed conversions, projected conversions, observed conversion value, projected conversion value, projected CPA, projected conversion rate, and projected ROAS. Recent periods show the largest gap between observed and projected — that gap is what Predictive Attribution fills. ### When projection matters most The biggest impact is on recent data for slow-converting campaigns. A prospecting campaign from last week might show 5 observed conversions — but the model projects 12 more will arrive. Without projection, you might cut the campaign. With it, you see the full picture and keep optimizing. For fast-converting campaigns like retargeting, the projection adds little — most conversions are already in. The system makes this distinction automatically, campaign by campaign. --- # Self-Reported Reattribution How SegmentStream makes invisible channels visible — LLM classification of free-text survey responses, stitched to sessions via identity graph and synthetic touchpoints. --- Attribution measures clicks. But entire categories of marketing — podcasts, TV, word of mouth, out-of-home — work through influence, not clicks. They are invisible to every attribution model. Self-reported reattribution makes them visible. ## 01 — The Visibility Gap Click-based attribution — as described in the [cross-channel attribution](/measurement-engine/cross-channel-attribution.md) methodology — tracks what it can see: ad clicks, organic search clicks, referral links. Every touchpoint in the journey needs a click with a URL parameter. If a channel does not produce a trackable click, it does not exist in your attribution data. Entire categories of marketing work this way. Podcasts, TV, word of mouth, out-of-home, influencer content, radio, events — these channels create awareness and drive action, but the action they drive is almost never a direct click. The user hears about you on a podcast, then searches your brand name. They see a billboard, then types your URL directly. Other channels fall in between. YouTube, TikTok, and AI chat (ChatGPT, Perplexity) generate some trackable clicks, but clicks capture only a fraction of their influence. A user watches a YouTube review, then searches your brand name a week later. They read a ChatGPT recommendation, then navigates directly to your site. The referral header exists, but the full influence is invisible to click-based attribution. Attribution records the brand search click, the direct visit, the organic landing. The channel that actually introduced the user is invisible. The credit flows to the last trackable touchpoint — which is almost always brand search, organic brand, or direct. > **[Interactive: SRA Comparison Table]** > Static table showing channel distribution before and after SRA correction. Nine base channels with columns for "Without SRA" and "With SRA" conversions plus percentage change. Direct/None drops from 847 to 339 conversions. Five new channels appear below a separator (Word of Mouth: 266, Podcast: 194, TV/Radio: 97, OOH: 73, Influencer: 73) — these had zero attribution before SRA. ### The absorption problem Brand search, organic brand, and direct traffic function as sponges. They absorb credit from every awareness channel that does not produce a click. A TV campaign that drives thousands of brand searches shows up as "Google Ads — Brand Search" in attribution. A podcast mention that sends listeners to your site shows up as "Direct / None." Word of mouth — your most valuable channel — is completely invisible. The distortion compounds in two directions. Awareness channels appear to produce zero return, making them impossible to justify in budget discussions. Meanwhile, brand search and direct show artificially high performance, absorbing credit they did not earn. Budget flows toward the channels that are easiest to measure, not the ones that are most effective. ### Why standalone SRA is not enough The obvious fix — ask the customer "How did you hear about us?" — is the right idea but incomplete on its own. Standalone self-reported attribution has structural limitations: - **Fragmented coverage** — even at 85-95% response rates, some users do not answer. The gap is not random — users in a hurry, mobile users, and returning customers skip the question more often. - **Response bias** — not all channels have equal response rates. Users who discovered you through memorable channels (a friend, a specific podcast) recall and report more reliably than those who saw a display ad they have already forgotten. - **Channel-level only** — a user can tell you "I heard about you on a podcast," but not which campaign or ad group drove the visit. SRA provides channel attribution, not campaign-level granularity. - **The triangulation fallacy** — a common response is to average SRA with click-based attribution, hoping the truth lies in the middle. It does not. Averaging two different measurement methods with different biases does not cancel the biases — it produces a third number with no clear meaning. Self-reported data is most valuable as a *correction layer* on top of click-based attribution — not as a replacement for it. The hybrid approach: attribution handles every channel that produces a trackable click. Self-reported data corrects the channels that attribution cannot see. On conflicts — attribution says Facebook, the user says search — the paid click wins. A verified click is stronger evidence than a recalled impression. SRA also measures something no click-based system can: true brand awareness. "Already knew about you," "word of mouth," and "a friend recommended you" quantify the organic strength of your brand — PR impact, customer referrals, and unprompted recognition. > Attribution sees clicks. SRA sees influence. Neither is complete alone. Combined as a hybrid — with attribution as the foundation and SRA as the correction layer — the invisible channels become visible without compromising the channels that are already well-tracked. ## 02 — Deploying the Survey One question. Free-text input. Placed as early as possible in the conversion flow. That is the entire deployment. ### Placement: earlier is better The survey should appear at the first moment the user provides identifying information — email capture, account registration, or checkout. Post-purchase is the most common placement and the worst. Response rates vary dramatically by placement: - **During registration or checkout** (optional field) — 85-95% response rate. The user is already filling out a form. One more optional field adds no friction. - **Post-purchase** — 50-60% response rate. The user has already completed their goal. A follow-up survey is an interruption, and response drops accordingly. Earlier placement also captures a broader audience. A post-purchase survey only captures buyers. A registration-time survey captures every user who creates an account — including those who never purchase. For businesses with long consideration cycles (B2B, high-ticket e-commerce), post-purchase misses the majority of the addressable audience. > **[Interactive: Survey Placement Comparison]** > Side-by-side comparison of two survey placements. Left panel "Registration / Checkout" shows a form with a "How did you hear about us?" field embedded naturally alongside name and email fields — labeled "User must complete this step." Right panel "Post-Purchase" shows a standalone survey card after order confirmation — labeled "User already got what they came for." ### Free-text over dropdowns Dropdowns are simpler to analyze but worse for measurement. Five reasons: 1. **Channel discovery.** A dropdown with pre-defined options cannot surface channels you did not think to include. Free-text reveals unexpected sources — specific podcast names, AI assistants (ChatGPT, Perplexity, Copilot), niche communities, and channels that did not exist when the dropdown was created. AI chat is showing meaningful volumes in production data. 2. **Retroactive reclassification.** Free-text responses can be re-classified at any time. When a new channel category emerges, you re-run the classifier on historical responses and retroactively surface the data. Dropdown responses are locked to the options available when the user answered. 3. **No priming bias.** A dropdown primes the user with options. If "Facebook" is listed first, users who are unsure will select it because it looks familiar. Free-text forces recall, which produces more accurate responses. 4. **Richer signal.** "My colleague Sarah mentioned you at a conference" is far more informative than a dropdown selection of "Word of mouth." Free-text captures context that structured responses cannot. 5. **Position bias.** Users disproportionately select the first option in a dropdown. Randomizing order helps but does not eliminate the effect. Free-text has no position bias. > **[Interactive: Free-Text vs Dropdown]** > Side-by-side comparison. Left: free-text input field capturing rich, unstructured responses. Right: dropdown selector spanning two columns, showing how pre-defined options limit discovery and introduce position bias. The cost of free-text is classification complexity — raw text needs to be normalized into channel groups before it can be used in attribution. This is the problem the next section solves. ## 03 — Classification and Mapping Collecting free-text responses is easy. Making them usable for attribution is the hard part. A single survey question generates 1,000+ unique text strings: "my friend told me," "heard on Joe Rogan," "saw an ad on Instagram," "google," "already knew about you." Each one needs to be mapped to a channel group that attribution can use. ### LLM classification An LLM classifier normalizes raw text into channel groups. The input is the user's free-text response. The output is a standardized channel label: "Word of Mouth," "Podcast," "TV," "AI Chat," "Out-of-Home." The classifier uses human-in-the-loop learning. Each human correction — "this response was classified as Social Media but should be Influencer" — enhances the classification prompt. Over time, the classifier converges on the project's specific channel vocabulary. Corrections are validatable via MCP: you can query the classifier to see every correction, every reclassification, and the current prompt. > **[Interactive: LLM Classification Flow]** > Three-column responsive layout. Left column: messy free-text responses in speech bubbles ("my friend told me", "saw on tiktok", "heard on podcast xyz", "google", "idk", "some ad somewhere"). Center: LLM classifier node with connecting lines. Right column: clean channel groups with triage categories color-coded — actionable (Friend/WOM, TikTok, Podcast), auto-ignore (idk), ambiguous (google, some ad somewhere). ### Triage: what to use, what to ignore Not every survey response carries attribution signal. The classifier sorts responses into three categories: - **Actionable** — maps to exactly one unambiguous channel. "My friend recommended you" maps to Word of Mouth. "Heard on a podcast" maps to Podcast. These responses override attribution where applicable. - **Auto-ignore** — carries no useful signal. Null responses, garbage input ("n/a," "asdf"), "Other," "Organic," and opt-outs are excluded from the override logic entirely. - **Ambiguous** — could mean multiple channels. "Social media" could be paid Meta, organic Instagram, or influencer content. These require data-driven resolution — checking which platforms have ad spend, which dominate — before they can be classified. ### The search engine problem "Google" and "search engine" deserve special treatment. When a user says "I found you through Google," they could mean brand search (paid), generic search (paid), or organic search. These are three fundamentally different channels with different strategic implications. No amount of spend data resolves this ambiguity — the question is *what* the user searched for, not *which* engine they used. The correct handling: ignore "search engine" responses for attribution override. Keep the original click-based attribution, which at least knows whether the click was paid or organic, brand or generic. Self-reported data adds nothing when click-based tracking already has the answer. ### Conflict resolution When attribution and survey data disagree, the rule is simple: paid clicks always win. If attribution recorded a paid Facebook click and the user says "search engine," the Facebook click stands. A verified click is stronger evidence than a recalled impression. Override logic applies to two categories of traffic only: - **Brand search and organic brand** — these channels absorb awareness credit. A user who heard about you on a podcast searches your brand name, clicks the paid ad or the organic result, and attribution credits the search. The SRA response reveals the true discovery channel. - **Non-paid traffic** — direct, organic, and referral sessions where no paid click exists. There is no paid evidence to protect, so the self-reported channel is the best available signal. Everything else — paid non-brand traffic where a verified click exists — keeps its original attribution. The override is a correction, not a replacement. This is the classification and mapping methodology the SegmentStream Measurement Engine implements. Responses are collected, classified by an LLM with human-in-the-loop learning, triaged for signal quality, and mapped to channel groups with conflict resolution rules that protect verified paid clicks. ## 04 — The Synthetic Touchpoint Classification determines *what* channel the user came from. The synthetic touchpoint determines *how* that information enters the attribution pipeline. The goal: make self-reported channels appear in attribution reports alongside click-tracked channels, with no special handling required downstream. ### How it works When a user's survey response produces an actionable channel and the override rules approve it, the pipeline creates a synthetic session record. This record looks exactly like a regular session in the attribution table — same schema, same fields — but with a fabricated timestamp: one second before the user's first recorded visit. First-click attribution picks up the synthetic touchpoint as the earliest session in the user's journey. The self-reported channel becomes the attributed source. No changes to the attribution logic, no special cases, no parallel reporting systems. The correction integrates directly into the existing pipeline. ### Identity graph bridging The survey is typically completed on one device — the one where the user registered or checked out. The first visit may have happened on a different device entirely. The [identity graph](/measurement-engine/identity-graph.md) bridges the gap: the user's `universal_id` links the survey device to the first-visit device. Without identity resolution, the synthetic touchpoint would have no journey to attach to. ### Override rules The synthetic touchpoint is only created when override conditions are met. Three types of traffic are treated differently: - **Brand search and organic brand** — overridden. These channels absorb awareness-channel credit by design. The self-reported response reveals the true discovery source. - **Direct / none** — overridden. No existing attribution signal to protect. The self-reported channel is the only available evidence. - **Paid non-brand** — protected. A verified paid click is stronger evidence than a survey response. The original attribution stands. > **[Illustration: Synthetic Touchpoint Before/After]** > Two horizontal timelines showing the same user journey. "Before SRA" timeline: three touchpoints (Direct/None visit, Organic search, Conversion) with first-touch attribution crediting Direct/None. "After SRA" timeline: a synthetic "SRA Signal" touchpoint is prepended one second before the first visit, attributing to TikTok. First-touch attribution now credits TikTok instead of Direct/None. ### The pipeline Five stages, each feeding the next: 1. **Survey collection** — free-text response captured at registration, checkout, or email capture 2. **LLM classifier** — raw text normalized to channel groups with human-in-the-loop learning 3. **Identity graph** — survey user linked to their full cross-device journey via `universal_id` 4. **Synthetic touchpoint** — session record created one second before first visit, with the self-reported channel as traffic source 5. **Attribution reports** — standard first-click reports now include self-reported channels alongside click-tracked channels > **[Illustration: SRA Pipeline Flow]** > Left-to-right process diagram with five stages connected by arrows: Survey (free-text input) → LLM Classifier (channel mapping) → Identity Graph (user matching) → Synthetic Touchpoint (session creation) → Attribution Reports (unified output). > The synthetic touchpoint is what makes SRA a correction layer instead of a parallel reporting system. Self-reported channels appear in the same reports, the same tables, and the same budget optimization logic as every click-tracked channel. One pipeline. One source of truth. ## 05 — What Goes Wrong and How to Avoid It 1. **Asking post-purchase instead of during registration.** Post-purchase surveys reach 50-60% of users. Registration-time placement reaches 85-95%. The gap is not just volume — post-purchase misses every user who creates an account but does not buy, which in B2B and high-ticket e-commerce is most of the audience. 2. **Using dropdowns instead of free-text.** Dropdowns are easier to analyze but cannot discover new channels, cannot be reclassified retroactively, and introduce priming and position bias. The classification problem is solvable. The missing-data problem is not. 3. **Overriding all traffic.** Only brand search and non-paid traffic should be overridden. Paid non-brand clicks are verified evidence — overriding them with survey responses destroys accurate data to replace it with less accurate data. 4. **Trusting "search engine" or "Google" at face value.** These responses are fundamentally ambiguous. They could mean brand search, generic paid, or organic. Click-based attribution already distinguishes between these — the survey response adds no information. Ignore them. 5. **Using raw text without classification.** Raw survey responses contain hundreds of variations for the same channel. Without LLM classification, you end up with a fragmented long-tail that is impossible to use for attribution. 6. **Treating SRA as standalone measurement.** Self-reported data has coverage gaps, response bias, and no campaign-level granularity. It is a correction layer for channels that attribution cannot see — not a replacement for click-based attribution. 7. **Skipping identity resolution.** The survey is completed on one device. The first visit may have happened on another. Without the identity graph bridging survey responses to first-visit sessions, the synthetic touchpoint has no journey to correct. 8. **Expecting campaign-level granularity.** A user can tell you "I heard about you on a podcast." They cannot tell you which ad group or bid strategy drove their awareness. SRA provides channel-level correction. Campaign-level optimization still depends on click-based attribution. ## 06 — See It Work Everything described above runs inside the SegmentStream MCP server. Four MCP methods expose the full SRA pipeline. The first returns the project's SRA configuration — extraction SQL, classifier settings, and override rules. The second shows raw survey answers with their LLM classification. The third aggregates by channel to show distribution after SRA correction. The fourth traces individual user journeys to show the before-and-after: original attribution versus corrected attribution. > **[Interactive: SRA Terminal Demo]** > Four-tab terminal demo in Claude Code style. Tab 1 "SRA Settings" runs `mcp__segmentstream__get_sra_settings` showing pipeline configuration and override rules. Tab 2 "SRA Answers" runs `mcp__segmentstream__get_sra_answers` showing raw survey responses with LLM classification results. Tab 3 "SRA Channels" runs `mcp__segmentstream__get_sra_channels` showing aggregated channel distribution with percentages. Tab 4 "SRA Overrides" runs `mcp__segmentstream__get_sra_overrides` showing before-and-after channel shifts for individual users. ### What the methods show `get_sra_settings` returns the project's pipeline configuration — whether classification is enabled, which classifier is active, and how override rules are structured. `get_sra_answers` shows every raw survey response and its LLM classification, making it easy to spot misclassifications and verify the triage logic. `get_sra_channels` aggregates by classified channel — which channels appeared and what share of conversions they account for. Channels that were previously invisible — Word of Mouth, Podcast, TV, Out-of-Home — appear alongside click-tracked channels with real conversion counts. Underreported channels like YouTube, TikTok, and AI Chat show their true scale. The "Direct / None" bucket shrinks as absorbed credit is redistributed to the channels that actually drove it. `get_sra_overrides` exposes the correction at the individual level: the original attributed channel, the self-reported channel, and the override decision. You can see exactly which users were reattributed, from which channel to which, and whether the override was triggered by the brand-search rule or the non-paid rule. ### Validating the output Every override is auditable. You can trace a reattribution to a specific `universal_id`, see the original survey response, the LLM classification, the override rule that fired, and the synthetic touchpoint that was created. The data lives in your BigQuery warehouse. The override logic is deterministic and configurable per project. ### What changes in practice Teams that deploy SRA correction typically see three shifts in their attribution data: - **Invisible channels appear** — Podcast, TV, Word of Mouth, Out-of-Home, and other channels with zero click tracking show up in reports for the first time with real conversion data. These were always driving value — attribution simply could not see them. - **Underreported channels show their true scale** — YouTube, TikTok, AI Chat, and prospecting social have some click tracking, but clicks capture only a fraction of their influence. SRA reveals the full contribution, typically 2-5x higher than click-only attribution. - **Brand search and direct shrink** — the credit sponges lose their artificially concentrated attribution as credit flows back to the channels that actually earned it. SRA also measures true brand awareness. "Already knew about you" and "word of mouth" responses quantify the channels that no click-based system can see: PR impact, organic brand strength, and customer referrals. --- # CRM Funnel Attribution Connect CRM, ERP, or any BigQuery table as a conversion source. Track every funnel stage from MQL to Closed/Won with proper attribution back to the original marketing click. --- Leads are fast but noisy. Revenue is true but slow. The channels that produce the cheapest leads are not the channels that produce revenue — and you cannot see the difference without connecting your CRM to attribution. ## 01 — Leads Arrive in Hours, Revenue Arrives in Months A Google Ads click happens in January. The lead fills out a form the same day. The deal closes in June. Five months separate the marketing event from the revenue event — and every decision made in between is based on incomplete data. Most marketing teams optimize on the fast signal: leads, form submissions, demo requests. These arrive within hours of the click. They are measurable, attributable, and available at campaign-level granularity. The problem is that lead volume does not predict revenue. The channels that produce the cheapest leads are often the channels that produce the worst pipeline. Display and broad-match search generate high lead volume at low cost. Sales works those leads for months and closes almost none. Meanwhile, a niche LinkedIn campaign produces fewer leads at higher CPL — but those leads close at 3x the rate with 2x the deal size. You cannot see this without connecting CRM data to attribution. Lead volume alone tells you which channels are cheapest. Revenue data tells you which channels are most profitable. These are different questions with different answers. ### Why you cannot just optimize on Closed/Won The obvious response: skip the leads, optimize directly on revenue. This does not work either. Closed/Won data has a 3–6 month lag. By the time you see which January campaigns produced revenue, it is July. The campaigns have changed, the bids have shifted, and the feedback is too stale for daily optimization. Closed/Won is also sparse. A B2B company might close 20–50 deals per month. That is not enough signal for campaign-level optimization — too few conversions to distinguish real performance differences from noise. > Leads are fast but noisy. Revenue is true but slow. You need both — and you need them connected to the same marketing touchpoints. CRM funnel attribution connects every stage from lead to revenue back to the original click. ## 02 — Every Stage from Lead to Revenue A B2B funnel is not one conversion. It is a sequence of stages, each with different timing, different volumes, and different signal quality: - **Lead** — form submission or inbound inquiry. Arrives within hours of the click. High volume, low signal. - **MQL** — Marketing Qualified Lead, meets a scoring threshold. Arrives within days. Moderate volume, moderate signal. - **HandRaiser** — the lead explicitly requested sales contact (demo request, pricing page). Immediate qualification. Lower volume, high signal. - **Opportunity** — enters the sales pipeline with a deal amount. Arrives weeks to months later. Low volume, high signal. - **Closed/Won** — signed deal. Arrives months later. Lowest volume, highest signal. Each stage is defined as a separate conversion in SegmentStream, each with its own custom SQL querying CRM tables in BigQuery. Every stage gets attributed independently — first-click, last significant click, and ML attribution run simultaneously across all stages. > **[Interactive: CRM Cohort Funnel Table]** > Static table showing a January cohort of leads tracked across the full B2B funnel over 6 months. Six channels (Google Ads Brand, Google Ads Generic, LinkedIn, Meta Prospecting, Organic Search, Word of Mouth via SRA) with columns for Leads, MQL, HandRaiser, Opportunity, Closed/Won, and Revenue. The highest value in each column is highlighted — Google Ads Brand leads at the Lead stage (142), but LinkedIn leads at Closed/Won (7, $520K revenue). Channel rankings flip visibly across stages. ### Channel rankings flip at every stage The cohort table reveals what lead-level reporting hides. Consider a January cohort of leads, viewed six months later across the full funnel. Google Ads — Brand Search leads the table at the Lead stage: highest volume, lowest cost. By Opportunity, LinkedIn has overtaken it. By Closed/Won, the ranking has flipped entirely. HandRaiser is the stage that separates signal from noise. Leads that explicitly request sales contact convert to Opportunity at roughly 57%, compared to roughly 38% for non-HandRaiser MQLs. A channel that produces fewer leads but a higher share of HandRaisers is almost certainly a better investment — you just cannot see it from leads alone. Multi-stage attribution makes these trade-offs visible. You can see which channels produce volume (leads), which produce intent (HandRaisers), and which produce revenue (Closed/Won) — all attributed to the same marketing touchpoints, in the same reports. ## 03 — Two Timestamps, One Conversion Every B2B attribution system faces the same tradeoff. A click happens in January. The deal closes in June. Where does attribution anchor its lookback window — on the click date or the close date? - **Anchor on close date (June)** — a standard 90-day window looks back to March. The January click is outside the window. Attribution fails silently — the deal appears unattributed, or worse, gets credited to whatever retargeting campaign touched the lead in April. - **Anchor on lead date (January)** — the 90-day window finds the click. But now reports show $50K of revenue in January that actually closed in June. The CFO sees phantom revenue. Monthly reporting breaks. Most systems force you to pick one and accept the distortion. Extend the window to 180 days and you catch more clicks — but misattribute them to the wrong time period. Shorten the window and you get clean timing — but miss the clicks that actually drove the pipeline. > **[Illustration: Dual-Date Timeline]** > Horizontal timeline from December to June. Three key events marked: December 5 ad click, January 10 lead created (`attribution_time`), June 15 deal closed (`conversion_time`). A 90-day lookback bracket anchored on January 10 extends back past December 5, catching the click. The dual timestamps are visually distinct — attribution time in accent color, conversion time in a secondary color — showing how they solve the window problem without compromise. ### Decoupling conversion time from attribution time SegmentStream separates the two timestamps that traditional systems conflate. Each funnel conversion has two dates: - **Conversion time** — when the business event happened. The deal closed on June 15. Reports, exports, and revenue dating use this timestamp. It matches the CRM and the P&L. - **Attribution time** — when to anchor the lookback window. The lead was created on January 10. The attribution engine uses this timestamp to find the marketing session that brought the lead — within a standard 90-day window. In the custom conversion SQL, the two dates map to two fields: - `created` (maps to `conversion_time`) — the close date, June 15 - `attribution_time` — the lead creation date, January 10 Result: the $50K Closed/Won appears in June reports (correct revenue timing), attributed to the Google Ads click from December (correct marketing credit). No window compromise. No phantom revenue. ### Automatic session lookback The pipeline reads the actual `attribution_time` values in each processing batch and extends its session loading range accordingly. If a June Closed/Won has a lead from January, the pipeline loads sessions from October (90 days before January) through June. A 2-week MQL and a 6-month enterprise deal are handled by the same system, with no manual configuration. ### Two report perspectives Reports expose both timestamps: - **By conversion time** — revenue appears when the deal closed. Matches the P&L. Use this for executive reporting and finance reconciliation. - **By click time** — revenue appears when the marketing click happened. Shows which month's campaigns produced revenue. Use this for campaign optimization and budget planning. Both perspectives come from the same underlying data. No duplication, no reconciliation needed. ## 04 — When CRM Data Changes After the Fact CRM data is not static. Deal amounts change. Stages revert. Records are updated weeks or months after creation. Attribution that treats CRM data as write-once will drift out of sync with reality. Three scenarios happen regularly: - **Value update** — a $50K deal is renegotiated to $75K. The conversion record needs the updated amount, and every report that includes this deal needs to reflect the new value. - **Stage reversion** — a Closed/Won deal reverts to Closed/Lost. The conversion should disappear from reports entirely — it is no longer revenue. - **Refund or credit** — post-close adjustments reduce the deal value to zero. Functionally identical to a reversion — the conversion value drops to zero and the record is removed. > **[Illustration: Adjustment Window Flow]** > Three horizontal lanes showing CRM change scenarios flowing through a MERGE operation to attribution re-run. Lane 1: "New record" (new lead qualified) → INSERT → attribution re-run. Lane 2: "Value changed" ($50K → $75K) → UPDATE → attribution re-run. Lane 3: "Deal reverted" (Won → Lost) → DELETE → attribution re-run. The UPDATE lane is highlighted in accent color. ### Adjustment windows Each funnel conversion has an adjustment window — the number of days the system re-queries CRM data to catch retroactive changes. Typical values: - **MQL / HandRaiser** — 14-day window. Stage qualification rarely changes after two weeks. - **Opportunity** — 30-day window. Deal amounts and stages shift during active negotiation. - **Closed/Won** — 180-day window. Enterprise deals can be renegotiated, clawed back, or credited months after close. When the adjustment window is active, the custom conversion SQL re-runs for the extended date range. A MERGE operation handles three cases: - **New record in source** — insert. A new deal was created or backdated. - **Value changed** — update. The deal amount or qualification flag changed. - **Record disappeared from source** — delete. The deal reverted, was disqualified, or was removed from the CRM. Attribution automatically re-runs for the affected date range. Reports always reflect the current CRM state — not a stale snapshot from the day the deal first appeared. ### Gross margin attribution Some businesses attribute gross margin instead of revenue. The conversion SQL returns margin as the `value` field instead of deal amount. The same adjustment window and MERGE pattern apply — when margins are recalculated, the updated value propagates through attribution automatically. ## 05 — The Ground Truth for Calibration CRM funnel data is not just for backward-looking reports. It is the strategic ground truth that calibrates the entire measurement engine. Teams that connect CRM funnel data to attribution typically see three shifts: - **Budget moves toward quality** — channels that produce expensive leads but high close rates gain budget. Channels that produce cheap leads with low close rates lose it. The reallocation is data-driven, not intuition. - **Quarterly reconciliation replaces monthly guessing** — Closed/Won data validates (or contradicts) the lead-based optimization decisions made months earlier. Teams stop arguing about channel quality and start measuring it. ## 06 — Connecting CRM Leads to Website Visitors Attribution needs a link between the CRM record and the website session that brought the lead. Without it, the deal is visible but unattributed — revenue exists in the CRM but no marketing channel gets credit. The [Identity Graph](/measurement-engine/identity-graph.md) handles this through deterministic matching across four signal types. Each signal has different coverage and reliability. CRM conversions use all four, layered from strongest to broadest. > **[Illustration: CRM Identity Stitching]** > Two cards — "CRM Record" (left, accent border) and "Website Sessions" (right) — connected by four horizontal lines of decreasing thickness representing match strength. From top to bottom: `anonymous_id` (direct cookie match, strongest), `email_hash` (SHA-256 across forms), `user_id` (CRM contact ID), `IP + timing` (3-day window fallback, weakest). Each signal is labeled with its mechanism. ### Anonymous ID The strongest signal. SegmentStream's SDK sets an `anonymous_id` cookie on every website visitor. When a lead fills out a form, a hidden field captures the `anonymous_id` and writes it to the CRM record. The custom conversion SQL passes it as `client_id`, and attribution links the deal directly to the session that produced it. This is a deterministic one-to-one match. No inference, no probability. The CRM record carries the exact cookie value from the session that created it. ### Email Hash When the CRM has the lead's email address, the system generates a SHA-256 hash and matches it against `email_hash` values in the identity graph. If the same email was submitted on the website — newsletter signup, checkout, account creation — the anonymous sessions from that email are linked to the CRM lead. Email hash also enables cross-device stitching. Marketing emails include the hash in URL parameters (`ehash`). When the recipient clicks on a different device, the SDK captures the hash and links the new session to the existing identity. ### User ID CRM systems that sync a contact or account identifier — Salesforce's `ConvertedContactId`, HubSpot's contact ID — can match against logged-in website sessions where the same `user_id` was set. This covers returning visitors who are already authenticated. ### IP and Timing When no direct identifier exists, the identity graph falls back to exact IP address matching within a bounded time window. If the CRM lead's IP address matches a website session's IP within a 3-day window, the sessions are linked. This is a deterministic match on an exact key with a time constraint — not a statistical guess. IP matching has known limitations. Shared office networks produce false positives. Mobile carrier IPs rotate. But as a fallback when stronger signals are unavailable, it recovers sessions that would otherwise go unmatched. ### The stitching order matters The identity graph runs Union-Find (connected components) across all signals daily. A single lead may connect to multiple anonymous sessions across devices — the form submission on desktop, the initial ad click on mobile, the email click on a tablet. Attribution sees the full cross-device journey, not just the session where the form was submitted. For implementation details on how the identity graph resolves these signals, see the [Identity Graph](/measurement-engine/identity-graph.md) page. ## 07 — See It Work Everything described above runs inside the SegmentStream MCP server. Two queries demonstrate the full pipeline: a Closed/Won attribution report by channel, and an individual user journey showing the gap between the original click and the revenue event. > **[Interactive: CRM Terminal Demo]** > Two-tab terminal demo in Claude Code style. Tab 1 "Closed/Won Attribution" runs `mcp__segmentstream__get_report_table` with conversion "Closed/Won", dimensions by channel, showing a table with columns Channel, Cost, Conversions, Revenue, ROAS. LinkedIn leads with 7 conversions and $520K revenue (18.3x ROAS). Insight: "LinkedIn produces 4x more revenue than Google Ads Brand despite 60% fewer leads." Tab 2 "User Journey" runs `mcp__segmentstream__get_user_journey` tracing a $75K LinkedIn deal from December 5 click through January lead creation, February MQL, March Opportunity, to May Closed/Won — a 5-month gap with dual-date attribution anchoring on January 10 to find the December click. Identity signals shown: `anonymous_id`, `email_hash`. ### Closed/Won attribution by channel `get_report_table` returns a standard attribution report filtered to the Closed/Won conversion. The output shows cost, conversions, revenue, and ROAS by channel — with attribution anchored on lead creation date, revenue dated to close date. You see which channels produced pipeline that actually closed, not just pipeline that entered. Compare this with the same report filtered to Lead or MQL. The channel rankings shift — often dramatically. Channels with the lowest cost-per-lead may have the highest cost-per-closed-deal, and vice versa. The multi-stage view makes these trade-offs visible in a single query. ### User journey: 5-month click-to-close gap `get_user_journey` traces an individual lead's path from first click to Closed/Won. The output shows every session, every touchpoint, and every funnel stage transition — with the identity signals that connected them. A typical B2B journey: Google Ads click in January (anonymous session), form submission in January (identity stitched via `anonymous_id`), MQL in February, Opportunity in March, Closed/Won in June at $50K. Five months, multiple devices, one attributed journey. The December click gets credit for the June revenue. ### What changes in practice Teams that connect CRM funnel data to attribution typically see three shifts: - **Budget moves toward quality** — channels that produce expensive leads but high close rates gain budget. Channels that produce cheap leads with low close rates lose it. The reallocation is data-driven, not intuition. - **Quarterly reconciliation replaces monthly guessing** — Closed/Won data validates (or contradicts) the lead-based optimization decisions made months earlier. Teams stop arguing about channel quality and start measuring it. --- [Marginal Analytics](/measurement-engine/marginal-analytics.md) tells you exactly where to shift budget — which campaigns are saturated, which have headroom, and what the optimal split looks like. The analysis is done. The recommendations are specific. But applying them is a different problem. Open Meta Ads. Adjust bids and budgets. Open Google Ads. Do it again. TikTok. Pinterest. By the time you've touched every platform, half the morning is gone. Next week, repeat. The worst part? You never know if last week's changes actually helped. No feedback loop. No way to connect the budget shift you made to the revenue change that followed. --- ## 02 — One-Click Execution SegmentStream takes the recommendations from Marginal Analytics and turns them into platform-ready changes. Shopping Campaign from $4,200 to $2,800. Prospecting from $2,500 to $3,600. Target ROAS adjustments where needed. Every change is specific, not directional. One click applies all changes directly to every ad platform. No spreadsheets, no logging into four dashboards, no copy-pasting numbers. > **[Illustration: Budget Hero]** > 9 campaign changes calculated and applied across 3 platforms in one click. After changes go live, SegmentStream tracks what actually happened — did revenue move the way the model predicted? This feedback loop tightens the curves and sharpens the next round of recommendations automatically. > **[Illustration: Optimization Overview]** > Cumulative gain, prediction accuracy, and adoption rate tracked week over week. --- ## 03 — Reinforced Learning Think of how a self-driving car learns. It doesn't just follow a static map — it drives, observes what happens, and adjusts. Every mile makes the model better. SegmentStream works the same way. Every week, the system predicts what a budget change will produce. After the change goes live, it measures what actually happened. The gap between prediction and reality feeds back into the model — automatically recalibrating the diminishing returns curves, adjusting for seasonality shifts, and correcting for creative fatigue. You don't need to retrain anything or tweak parameters. The system improves on its own. Week 1 recommendations are good. Week 10 recommendations are sharp. Week 30 recommendations know your account better than any analyst could. This is the difference between a static optimization tool and an autonomous one. Static tools give you the same quality of output whether you've used them for a day or a year. A reinforced learning loop compounds — every cycle of predict, apply, measure makes the next cycle more accurate. > **[Illustration: Feedback Loop]** > Each cycle of predict, apply, measure, learn improves prediction accuracy from 89% to 98%. --- ## 04 — Best Practices Automated execution delivers the best results when you set it up right. Here are the most common pitfalls. ### Going aggressive on day one Cutting a $30K channel to $8K in one week resets platform delivery algorithms. First-week performance will be worse than the model predicts. **Fix:** Start with conservative scenarios. Increase aggression over 2–3 weeks as results validate. ### Ignoring platform learning phases Cutting Meta from $25K to $12K forces its algorithm to re-learn delivery. The transition dip is real, not a model error. **Fix:** Use platform-aware pacing that accounts for learning periods, or apply changes gradually. ### Treating optimization as a one-time event Curves shift with seasonality, creative fatigue, and competitive pressure. A quarterly reallocation is barely better than guessing. **Fix:** Run weekly. The system recalculates automatically — the only cost is reviewing and approving. --- # Incrementality Testing > Geo holdout experiments with synthetic control groups that measure what your ads actually add — not what attribution claims they do. Are your ads driving sales — or just taking credit for them? Find out with SegmentStream GeoLift Experiments. *8 min read · March 2026* ## 01 — The Case for Incrementality Testing [Cross-channel attribution](/measurement-engine/cross-channel-attribution.md) is the foundation of daily optimization. It connects clicks to conversions, provides granular performance insights, and feeds platform bidding algorithms. But attribution has blind spots. Take Google Brand Search. You bid on your own brand name. In analytics, the campaign shows outstanding ROAS. But would those conversions have happened anyway? If someone searches your brand name, they already know you. Without the paid ad, they'd likely click the organic link and convert regardless. You might not lose a single sales dollar — while saving a huge portion of your paid media budget. This is exactly why incrementality tests were invented — to measure what ads actually add to your overall sales. --- ## 02 — Two Types of Incrementality Tests ### In-Platform Lift Studies (and Why to Avoid Them) Meta, Google, TikTok, and Snap offer built-in "lift studies" that randomly split users into test and control groups. Two problems make them unreliable: **Unequal tracking.** Test groups get pixels, click IDs, and client-side identifiers. Control groups rely on server-side CAPI only. Test groups show more observed conversions because measurement is better — not necessarily because ads worked. **Blackbox.** You can't see how audiences were split, how conversions were counted, or what methodology was used. The platform selling ads is also grading them. > The platform selling you ads is also measuring whether those ads work — using methods you cannot audit. Avoid in-platform lift studies at all costs. ### Geo Holdout Experiments Geo holdouts split at the regional level — countries, states, cities, DMAs, or ZIP codes. Ads run normally in control regions and are paused in holdout regions. Geography is a clean boundary: no cross-contamination, no audience overlap, and you control the entire experiment. Instead of analyzing performance on a user level, you look at total sales across regions — without any attribution at all — and check: did sales go up or down because certain ad activity was turned off? This is the most accurate methodology for measuring incrementality — and it's exactly what SegmentStream offers. --- ## 03 — SegmentStream's Approach to Rigorous Geo Holdout Testing > **[Illustration: Experiment Phases]** > End-to-end experiment flow showing the sequential phases of a geo holdout test. ### Market Selection SegmentStream analyzes historical sales data across all available regions — factoring in seasonality, trends, and patterns. This surfaces regions that historically behave similarly, so they can be used for a valid experiment. Once the correlated regions are identified, SegmentStream splits them into two groups — test and control — so that each group's aggregate sales track in parallel: same peaks, same dips, same growth rate. Pause ads across the test group, keep them running across the control group, and you have a clean comparison. > **[Illustration: Region Correlation]** > The system identifies regions with correlated sales trends, then assigns them to test and control groups. ### Synthetic Controls Even correlated regions almost never have the same absolute sales volume. California might do $800K/month while Ohio does $420K. Comparing raw numbers would be meaningless. Synthetic control modeling solves this by applying weighted coefficients to each region, scaling them to a common baseline. The result: an apples-to-apples comparison where the only variable is whether ads were running or not. > **[Illustration: Synthetic Control]** > Raw sales at different absolute levels, then weighted to a common scale for valid comparison. ### Minimum Detectable Effect Every experiment has an MDE — the smallest lift it can reliably detect. SegmentStream calculates MDE upfront based on the number of regions, volume per region, and baseline variance. If the expected effect is smaller than MDE, the system flags it before the test runs — preventing wasted experiments that were never powered to produce a result. > **[Illustration: MDE]** > An 18% lift clears the 15% MDE threshold. A 5% lift is invisible — the test can't detect it. ### A/A Validation Before pausing any ads, SegmentStream runs an A/A period — both groups receive identical treatment. If the groups track closely, the experiment design is confirmed valid. If they diverge, the system flags the issue and adjusts region selection automatically. This catches competitor launches, seasonal anomalies, and assignment bias before they contaminate results. > **[Illustration: A/A Validation]** > Left: groups track together — valid design. Right: groups diverge — fix selection before proceeding. ### Sales Cycles SegmentStream accounts for your sales cycle length when designing the experiment. If your average cycle is 14 days, conversions at the end of the holdout were influenced by ads that ran before it started. The system automatically extends the test window or excludes the contaminated tail from evaluation. > **[Illustration: Sales Cycle Lag]** > The lag zone at the end of the holdout is contaminated by pre-test ad exposure. ### Evaluation Once the holdout period ends, SegmentStream doesn't just compare raw sales numbers between test and control groups. It applies the same modeling coefficients that were used to create synthetic controls — scaling sales from each region up or down to ensure incrementality is calculated on a normalized basis. This is a common mistake when running geo tests manually: marketers look at raw sales in their analytics reporting, see different numbers between regions, and draw conclusions — completely ignoring the synthetic control modeling that makes the comparison valid in the first place. > **[Illustration: Geo Experiment Chart]** > Control and holdout track closely pre-test. During the experiment, holdout drops. The gap is incremental lift. But a single incrementality number by itself is meaningless without knowing how confident you can be in it. That's why SegmentStream always reports results with **confidence intervals** — the range within which the true effect most likely falls. If a test shows +35% incremental lift with a confidence interval of [22%–49%], the entire range is above zero — the result is statistically significant. But if another test shows +5% with an interval of [−8%–+18%], the range includes zero, meaning the result is inconclusive. The ads might be incremental, or the observed difference might be noise. > **[Illustration: Confidence Interval]** > Google brand search: significant (CI above zero). TikTok awareness: inconclusive (CI spans zero). --- ## 04 — Geo Lift Experiments Made Easy Before, running or analyzing an incrementality test wasn't possible without a data science expert. With SegmentStream, you can design, launch, and evaluate incrementality tests — without any technical skills, knowing you will get trustworthy insights. Manage Geo Holdout Experiments from the SegmentStream UI: Or directly from your favourite AI tools like Claude Code, Claude Cowork, Cursor, or Codex: > **[Interactive: Incrementality Cowork Demo]** > AI conversation demo showing how to design and evaluate an incrementality test through natural language. --- ## 05 — Important Considerations Incrementality testing often sounds great in theory, but in practice — most tests are done wrong, or shouldn't be launched in the first place. Here are the most important things to understand before running one. ### Geo Holdouts Can't Measure Long-Term Brand Effects Many teams think that if last-click attribution can't measure upper-funnel, they need to jump straight to incrementality tests. The problem: most geo holdouts run for 2–4 weeks. If your goal is to measure how upper-funnel ads influence brand recognition and organic demand over time, a short holdout is simply the wrong methodology. Brand effects compound over months — a 4-week test won't capture them. ### Geo Holdouts Have Hidden Costs The primary cost isn't the platform fee or the setup effort. Tests are invasive — you stop showing ads to a large chunk of your audience. That means a noticeable drop in sales during the test period. You should only run an incrementality test if *not knowing* incrementality could lead to even bigger losses over time. ### Incrementality Means Nothing Without Confidence Intervals A headline number like "+12% lift" is meaningless if the confidence interval is [−5%, +29%]. The range includes zero, so you can't tell whether the ads had any effect at all. Always look at the interval, not the point estimate. ### Forgetting to Adjust Budgets in the Control Market When you remove test regions from targeting, the freed budget overflows into control regions — making them spend more than they should. This breaks the entire purpose of the test. It's one of the single biggest reasons why geo holdouts produce unreliable results. > At SegmentStream, we take incrementality testing seriously — and would never recommend running a test unless the specific use case requires it and there is no substitute methodology. But when a test makes sense to run, you can be confident the results are ones you can actually trust. --- # Server-Side Conversion Tracking Stop letting ad-platform pixels broadcast your audience to your competitors. Send only the conversions that matter, with the identity signals that make cross-device matching real. *8 min read · May 2026* > **[Visual: Pixel Broadcast Hero]** > A stylised website at the top fans red outbound arrows to eight ad-platform tiles (Meta, Google, TikTok, Pinterest, Snap, X, Amazon, The Trade Desk). Lighter dashed arrows then leak from the bottom-row platforms down to five competitor badges labelled "buying audience." The reader takeaway — every in-page pixel becomes a redistribution channel that pipes your audience to direct competitors via the open RTB graph. --- ## 01 — Your Pixels Are Broadcasting Your Audience to Your Competitors When an ad platform places a pixel on your website, that pixel sees every visitor who lands on the site — not just the visitors the platform itself drove. The platform uses those observations to learn what your audience looks like: their browsing patterns, demographic signals, interest categories, time-of-day behaviour. That learning feeds the platform's ad-targeting models. There are two consequences for the brand. - **The platform sells your audience to other advertisers.** The interest signals it learns from your traffic are available to anyone buying ads on that platform, including direct competitors. They can target lookalikes of your buyers, or people who showed interest in your category, without ever having visited your site. - **Your own ROAS becomes harder to measure honestly.** The platform attributes conversions to itself based on every pixel hit, so its self-reported ROAS counts brand-search, organic, and direct-traffic conversions that would have happened anyway. The pixel's optimisation gets credit it did not earn. When the platform also performs cookie-syncing — exchanging your visitor's identity with other ad networks — the audience leaves that single platform and enters the open RTB ecosystem. From there, dozens of demand-side platforms can buy access to your audience, including those used by your direct competitors. > One audit we ran: twelve ad platforms had pixels firing on the site. Two of them, Meta and Google, delivered any meaningful first-click revenue. The other ten observed every visitor and returned nothing. The Trade Desk pixel was actively broadcasting the audience into the open RTB graph for any DSP buyer to bid for. The pixel is the platform's primary monetisation surface. Every visit it captures becomes training data the platform resells. Strong brands pay the highest tax: the bigger your audience, the more competitors are willing to pay to retarget against it. Server-side conversion tracking is the structural fix. Instead of letting the platforms watch every visit, you tell each platform only about the conversions that matter, server-to-server, with the identity signals that make those conversions match correctly. --- ## 02 — What Server-Side Conversion Tracking Is In the in-page-pixel model, every visit to your site triggers a request from the visitor's browser to the ad platform. The platform receives a continuous stream of pageviews, add-to-carts, scrolls, and conversions, all stamped with that platform's cookies. The platform sees everything. In the server-side model, the visitor's browser hits a first-party endpoint — your own backend, or a measurement layer like SegmentStream. Events accumulate on your side. You decide which events qualify as conversions, attach the identifiers that carry the most matching power, and forward only those events to each ad platform via its server-side API. The platform sees only what you choose to send. > **[Visual: Pixel vs Server-Side]** > Two-panel architecture comparison. The left panel ("In-Page Pixel") shows the visitor browser hopping through an ad-platform pixel that loads on every route, then to the ad platform itself, with a high "audience exposure" bar. The right panel ("Server-Side") shows the same visitor hopping into a first-party endpoint on your domain that qualifies events server-side, then a dashed hop to an ad-platform CAPI that receives only qualifying conversions, with a low "audience exposure" bar. The reader takeaway — same visitor, same platform, but the data the platform gets to keep collapses from "every page view" to "only converting events." The shift changes three things at once. - **Scope.** The ad platform stops seeing pageviews, browse behaviour, and abandoned carts. It sees only converting events. Audience-leakage drops to whatever you choose to forward. - **Identity.** Server-side payloads can carry hashed first-party identifiers — email, phone, customer ID — which match across devices in ways third-party cookies cannot. Cross-device match rates improve. - **Attribution.** Because you control which events get forwarded, you can filter through an incrementality lens before sending: forward conversions that the platform actually drove, rather than every conversion the pixel happens to witness. The platforms have built their own server-side APIs precisely because privacy regulation, ad-blocker adoption, and ITP cookie restrictions have made the in-page pixel less reliable every year. Meta's Conversions API, Google's Enhanced Conversions, TikTok Events API, Pinterest Conversions API, Snapchat CAPI, LinkedIn CAPI — every major platform now exposes a server-side endpoint. Server-side conversion tracking is moving from optional optimisation to the default delivery mechanism. --- ## 03 — How the Architecture Works A working server-side conversion-tracking stack has four stages. Each stage answers a specific question; together they replace the in-page pixel without losing the conversion signal the ad platforms need to optimise. > **[Visual: Server-Side Flow]** > A horizontal four-step pipeline — Collection (first-party endpoint, with site events, server-side instrumentation, CRM sync) → Identity (unified user, click_id + ip, email hash, cross-device stitch) → Qualify (drop pageview / browse, drop AtC, keep purchase / lead) → Dispatch (Meta CAPI, Google Enhanced, TikTok Events). A footer note reads "Hashed identifiers only · no raw PII leaves your domain." The reader takeaway — server-side tracking is a pipeline, not a single switch, and each stage filters or enriches the signal before it leaves your infrastructure. ### Stage 1: Collection Events are captured from two sources. The website SDK reports user behaviour and click identifiers (gclid, fbclid, ttclid, li_fat_id, msclkid, and the rest) to a first-party endpoint. Server-side instrumentation reports conversions from the place they happen: the order-confirmation handler, the CRM stage transition, the lead-qualified webhook. Both streams land in the same warehouse. ### Stage 2: Identity stitching The collection layer cannot tell that a click on a mobile session and a purchase on a desktop session belong to the same person. The [Identity Graph](https://segmentstream.com/measurement-engine/identity-graph.md) does. It deterministically stitches anonymous events to known customers using the strongest available combination of identifiers: click ID, IP address, browser fingerprint, login event, hashed email match. Stitched users are the unit every downstream stage operates on. ### Stage 3: Qualification Most of what the in-page pixel sent to the platform was noise: pageviews, add-to-carts, scrolls. The qualification stage keeps only events that match a real conversion definition. That definition can come from your [CRM funnel](https://segmentstream.com/measurement-engine/crm-funnel-attribution.md) (Lead, MQL, SQL, Closed/Won), from purchase events with non-trivial revenue, or from a custom predicate. Anything below the threshold stays in your warehouse and never reaches the ad platform. ### Stage 4: Dispatch Qualifying events are forwarded to each ad platform's server-side endpoint with the identifiers that platform's advanced-matching scheme expects. Different platforms accept different subsets — Meta CAPI takes hashed email, phone, name, address, IP, user-agent, fbp/fbc click cookies, plus external_id; Google Enhanced Conversions takes hashed email and phone with the user-data fields; Pinterest, Snapchat, and TikTok have their own near-overlapping schemas. The dispatcher normalises identifiers (lowercase email, E.164 phone, country code prefix) and SHA-256 hashes them before they leave the server. By the time an event reaches Meta or Google, it carries the converting customer's identity in a form the platform can match against its own user graph — without exposing any PII to the browser, the network, or anyone watching client traffic. --- ## 04 — Cross-Device Advanced Matching with CRM Data The single biggest reason to operate server-side is cross-device match rate. A growing share of buyer journeys starts on a mobile ad and ends on a desktop browser hours or days later. Third-party cookies cannot follow that journey. First-party identifiers from your CRM can. Advanced matching is the formal name for the technique. When a server-side conversion event reaches an ad platform, it can include a small set of identifiers that uniquely describe the converting customer: hashed email, hashed phone, hashed first and last name, hashed city and ZIP, IP address and user-agent, plus your internal customer ID and the platform's click cookies (Meta's fbp/fbc, Google's gclid). The platform takes those identifiers and matches them against its own logged-in user graph. If the same email or phone exists on a Meta or Google account that clicked an ad three days earlier, the conversion gets credited to that click — even if the click was on a different device, browser, or cookie. > **[Visual: Cross-Device Match]** > Three vertically-stacked moments connected by a hashed-email bridge. Day 1: a mobile Instagram click captures `fbclid_AbC123…` anonymously. Day 6: the same person signs in on desktop, the CRM enriches the conversion event with `user@brand.com`. +ms later: the server-side dispatcher forwards the conversion to Meta CAPI carrying `sha256(email)` and the original click_id. A green confirmation card reads "Meta credits the Instagram click — hashed email matches the signed-in user on Meta's side." The reader takeaway — the bridge that connects a mobile click on Tuesday to a desktop conversion on Friday is a hashed CRM identifier, not a cookie. ### Why CRM data is the unlock Most brands already have the identifiers advanced matching needs — they live in the CRM. Email and phone are captured at sign-up. First and last name are captured at checkout. Customer ID is the primary key. The reason these identifiers are not reaching ad platforms today is not that the data is missing; it is that the in-page pixel cannot transmit them safely. PII cannot ride the cookie wire without exposing every customer to anyone watching that customer's browser. Server-side delivery is the only safe route. Once CRM identifiers are in the conversion stream, two measurable things happen. - **Match rate climbs.** Meta and Google publicly quote 5–15% incremental conversion attribution lift from advanced matching, on top of base server-side delivery. In categories where buyers research on mobile and convert on desktop, the lift can be considerably higher. - **The platform algorithm learns from a cleaner signal.** When more of the actual converters can be matched back to their original click, the platform's optimisation model gets a more accurate picture of which clicks lead to revenue. Smart bidding becomes smarter. ### The role of the Identity Graph Advanced matching only works when the right hashed identifier is attached to each conversion event. That requires knowing, at the moment a conversion happens, which CRM record the converting session belongs to. The [Identity Graph](https://segmentstream.com/measurement-engine/identity-graph.md) is the joinable layer that makes that resolution deterministic. Without it, you would either over-send (every event with every possible identifier, which is wasteful and privacy-loose) or under-match (relying only on the click cookie, which is exactly the cookie advanced matching exists to supplement). The same Identity Graph that powers cross-channel attribution powers advanced matching. The infrastructure investment compounds. --- ## 05 — What SegmentStream Provides Server-side tracking is a stack, not a single product. We are deliberate about which parts we own and which parts you keep close to the ad-platform contract. > **[Visual: SegmentStream Role]** > Three stacked ownership cards. Top (indigo, "SegmentStream") — Audience-Leakage Audit, Identity Graph, CRM Funnel Attribution, Recommendations. Middle (neutral, "You implement") — Meta CAPI, Google Enhanced Conversions, TikTok / Pinterest / Snap Events. Bottom (muted, "Stays in-page") — GA4 (aggregate analytics, not ad optimisation), Session recording (Clarity / Hotjar with identity sync OFF). The reader takeaway — SegmentStream owns the audit, identity, and attribution layer; the customer keeps each platform-native dispatcher under their own ad-account contract; analytics and replay tools stay in-page with cross-platform identity sync disabled. ### The audit The **audience-leakage audit** is a SegmentStream Agents skill that measures the gap on your site, per platform. It pairs a pixel inventory with first-click attribution to quantify how much of your audience each platform observes versus how much revenue that platform actually returns. The output is a per-platform tier list and a remove-first recommendation. Run this before any architectural change so you know which platforms to retire entirely, which to keep via server-side, and which to reduce in scope. ### The identity layer The [Identity Graph](https://segmentstream.com/measurement-engine/identity-graph.md) stitches anonymous and identified events into a single customer record using deterministic keys (click ID, IP, hashed email, login event, browser fingerprint). [CRM Funnel Attribution](https://segmentstream.com/measurement-engine/crm-funnel-attribution.md) joins your CRM as a first-class conversion source. Together they are the layer that makes advanced matching possible: at the moment a conversion happens, the right hashed identifier is already attached to the event. ### The dispatcher (your side) Each ad platform's server-side endpoint is best owned by you, not by us. Meta CAPI, Google Enhanced Conversions, Pinterest CAPI, TikTok Events API, Snapchat CAPI, LinkedIn CAPI — each is a contractual relationship between your ad account and the platform. The most resilient pattern is a warehouse-to-CAPI pipeline that reads qualifying events from the warehouse SegmentStream populates and forwards them to each platform with the identifier subset that platform accepts. That keeps the platform contract in your control, keeps the dispatcher auditable, and keeps you free to swap platforms without re-platforming the measurement layer. ### What stays in-page Server-side delivery does not eliminate every in-page tag. Your own product analytics (GA4 with measurement-protocol calls, Amplitude, Mixpanel) and session-replay tools (Clarity, Hotjar) remain valuable for product research and optimisation. What you remove is their cross-platform identity sync — the part that ports your audience to the ad-platform side of the same vendor. Clarity, in particular, exposes a project-level toggle for this. Server-side conversion tracking changes who sees what. Pixels stop watching every visit; ad platforms receive only the converting events you choose to send, with the cross-device identity that makes those events match correctly. The audience stops being broadcast. Long-term ROAS recovers. --- # Frequently Asked Questions > Answers to questions that come up in conversations with prospects and customers and aren't covered on other pages. For how engagements are scoped and priced, see [Pricing](https://segmentstream.com/pricing.md). ## Can I send my own conversion values, or do I have to use SegmentStream's ML models? You can use your own. SegmentStream lets us create custom conversions for your project that can carry any value you compute — including the output of your own pricing or scoring engine. That value is then used everywhere: attribution reports, budget optimization, and conversion exports to ad platforms like Meta and Google. ## Can I exclude certain orders from conversions — free samples, $0 products, test orders? Yes. Your purchase conversion doesn't have to include every order blindly. We can define it to exclude orders by value, SKU, product type, or any other order parameter — so samples and test orders don't pollute conversion counts, CPA, conversion rate, or revenue. If useful, the excluded orders can still be tracked separately as their own conversion. ## Can SegmentStream separate new-customer revenue from returning-customer revenue? Yes. We can create a separate conversion for first-time customer purchases, based on your store or CRM customer data, and use it as the primary view for evaluating acquisition campaigns. A second conversion for all purchases keeps total revenue visible — and reports compare both at channel and campaign level: one view for acquisition quality, one for total business impact. ## Is BigQuery a hard requirement? I don't have a data warehouse. No. Using your own BigQuery is completely optional. There are two options: SegmentStream-hosted storage — a managed data warehouse we provision and run under the hood (isolated dataset, EU or US region, your data stays yours and is exportable at any time) — or self-hosted, where you bring your own BigQuery project. You don't need to add a warehouse on your side. ## Meta restricts conversion tracking for my category (Health & Wellness, etc.). Can SegmentStream still optimize my Meta campaigns? Yes. Meta doesn't allow advertisers in restricted categories to send lower-funnel conversion events like leads or purchases for campaign optimization, even server-side. SegmentStream works around this with behaviour-based Synthetic Conversion Events: instead of reporting the actual conversion, Visit Scoring evaluates on-site behaviour and generates policy-safe events with predicted values — no restricted terminology in event names, parameters, or URLs. These events are exported to Meta via the Conversions API and give Meta's bidding a value signal to optimize toward, without ever seeing a restricted event. Note: value-based optimization on custom events is a Meta beta — ask your Meta representative to activate it on your account. ## Is SegmentStream HIPAA compliant? Do you sign a BAA? SegmentStream is not a HIPAA business associate and does not sign BAAs. The platform is built for marketing measurement on public websites and ad platforms — it is not designed to receive or process protected health information, and you should not send PHI into it. Most marketing and referral businesses in health-adjacent verticals are not HIPAA covered entities; the rules that usually apply to them instead (the FTC Act, the FTC Health Breach Notification Rule, and state laws like Washington's My Health My Data Act) hinge on user consent for sharing data with ad platforms. SegmentStream supports consent-aware tracking and data minimization, but the consent setup on your site is yours to own — and for any regulated vertical, have your counsel review the configuration. We provide a DPA and full security documentation. ## My sales cycle is long — lead values change as deals progress. Can conversion values be updated after the fact? Yes, in two complementary ways. Conversions adjustment re-syncs conversion data from your CRM within a configurable window, so when a lead's status or value changes weeks later, SegmentStream's reports and optimization reflect it. And for staged funnels (lead → qualified → closed), each stage can be tracked as its own conversion that inherits attribution from the original lead — so revenue closing months after the first click is still credited to the channel that generated it, without needing a months-long attribution window. One caveat for ad-platform exports: Meta only accepts events up to 7 days old, so late value corrections update your measurement and budget decisions, not events already exported to Meta. ## How can I verify that SegmentStream's numbers are right? Everything is designed to be traceable and auditable. For tracked conversions, you can open any customer journey and see the actual touchpoints behind the attribution — which campaign introduced the user, what happened later, and why credit was assigned the way it was. Predictions are back-testable: the model forecasts how an immature cohort will mature, and you compare the forecast with reality once enough time has passed. ## I already run server-side tracking — does that conflict with SegmentStream? No. Server-side tracking is about how events reach your analytics platform and ad platforms. SegmentStream consumes the data downstream: events come from your analytics system, and spend and campaign data comes from ad-platform APIs — so everything works the same whether your events arrive client-side or server-side. ## How long does implementation take? Connecting website tracking and ad platforms usually takes no more than a day. Time to first meaningful numbers depends on historical data: if it's already available in your data warehouse, reports populate right away; if data collection starts from scratch, expect 7–14 days. CRM integration takes a few days at most when access is granted and a native connector exists, and up to a week for a custom integration.