Customer segmentation means grouping your customers by what they have in common, so you can treat different groups differently. Behavioral segmentation groups them by what they do, purchases, browsing, engagement, which matters because a customer's last ten actions predict their next one better than their age bracket does. Most marketing tools in 2026 claim to do this, as does an entire industry of platforms. However in 2026 this market just went through its biggest shake-up in years, and the tools in it now disagree about the definition of segmentation.

Databricks launched CustomerLake in June, building the segmentation-and-activation layer straight into the data platform, which challenges standalone CDP vendors the same way data platforms adding analytics put on BI tools. The CDP Institute reports warehouse-native vendors growing headcount at roughly six times the industry average while the packaged incumbents flatline. The pattern is the same across the whole data market, and we've written about that bigger split separately: the standalone middle layer is getting absorbed from below.
Our CPO Ian Liddicoat has spent years building segmentation models for brands across retail, travel and financial services, so you'll get his take through this piece rather than just technical details.
Most tools treat a segment as a list: a set of customers matching rules, built to receive a campaign. A smaller group treats a segment as a model: a statistically discovered group with a value attached, built to feed every other piece of analytics you run. Which one you need decides which camp to shop in.
This camp: you already have the tool, and it already has your customers, so do segmentation there.
Klaviyo is the best version for pure ecommerce. It works on your Shopify data, keeps behavioral segments updated in real time, and layers AI on top for predicted lifetime value and churn risk. If you're a B2C brand running email and SMS, it's hard to beat for the price. Its issue is its scope: Klaviyo segments only what Klaviyo can see, and activation stops at its own channels plus the big ad platforms. HubSpot does the same job for B2B contact lists inside the CRM, and MoEngage and Insider do a similar job for mobile-first and cross-channel lifecycle brands respectively.
They all have the same issue though. Each of these tools segments its own copy of the customer. Run three of them and you get three different segments for the same customer. For a small brand on one channel that's a non-problem. It becomes a problem when your segmentation is supposed to inform anything beyond the next campaign.
This camp: the fix for fragmented customer data is a dedicated platform that unifies it, then segments from that unified profile.
Twilio Segment is the dev-friendly version, with a huge connector catalog for collecting events and pushing audiences out. Salesforce Data Cloud is the biggest player in this camp, effectively the default when Salesforce is already your CRM, now with Agentforce attached for AI-driven activation. Adobe Real-Time CDP plays the same role in Adobe's ecosystem, and ActionIQ sells enterprise governance to regulated industries.
These platforms solve the three-versions-of-the-customer problem, and for large B2C enterprises with real-time personalization needs, they’re a good choice. There are two drawbacks to this group though. First is cost and weight: enterprise licensing commonly runs six figures a year prior to implementation, and industry research puts CDP project failure rates at 30 to 40 percent, mostly on data readiness, which is a polite way of saying the platform arrived before the data foundation did. The second is that packaged CDPs keep a separate copy of your customer data, and the market is moving away from paying to duplicate data you already store.
The newest architecture, and the one with the most momentum: leave the data in the warehouse you already run, and do segmentation there.
Hightouch builds audiences directly on Snowflake, BigQuery, or Databricks and syncs them to your marketing tools, no second copy of the data, with AI features layered on. Census, now owned by Fivetran, and RudderStack run the same warehouse-native play with different focuses. For companies that already have a well-modeled warehouse and a data team keeping it clean, this camp is the new default, and it's where the analyst momentum clearly points.
The drawback is that warehouse-native tools assume the warehouse is in a good state. If your customer data isn't already consolidated and modeled to the point people trust it, there's nothing for these tools to segment, and the vendors themselves will tell you the prerequisite is a mature data function. They've also just inherited a strategic problem, because Databricks building CustomerLake into the platform means the layer they sell is being absorbed by the platform they depend on.
Everything above produces lists, just better or worse governed. The fourth camp comes at segmentation from analytics instead of marketing.
Here's how Ian describes what a segmentation model actually is, from his piece on AI customer segmentation: you're grouping customers who share behaviors and habits "with a complete understanding of their value to the business," and what you end up with is 20 to 30 named segments for a typical B2C business, each with a label, a profile, and a description built from the underlying data. The technique underneath is K-means or hierarchical clustering in Python, and as Ian points out, the statistics haven't changed much. "A lot of the equations have been around for a long time. What's changed is compute horsepower, the availability of data, and the ability to put large language models on top of analytics."
The part most segmentation shopping misses is what the segments are for. In Ian's words: "Segmentation isn't a standalone exercise. It's a component that shows up in your media mix model, your attribution model, your pricing model, because you want to understand the relative behavior of a given segment at any point." A price elasticity model responds differently by segment. A media mix model allocates differently by segment. That's the list-versus-model split in practice: a list can receive a campaign, but only a model can tell you which segment's demand is price-driven and which is media-driven, which is the question the CEO is actually asking.
Kleene is our version of this camp. The segmentation model is built on the consolidated platform (so the foundation problem that sinks 30 to 40 percent of CDP projects is handled as part of the engagement, not assumed), value and lifetime value are intrinsic to the segments, the segment labels feed the other KAI models, and KAI's natural language layer means you can ask about a segment's characteristics, what messaging it responds to, and through which channels, in plain English.
You can also go do it yourself: a capable data science team with Python and the same clustering techniques can build real segmentation models in-house, and for companies with that team, it's a possible path. The trade-off is the one Ian names about why an embedded partner moves faster: "If we're doing demand forecasting and we already have access to their warehouse, we probably built their third-party connectors. We know a huge amount about their data, their business, and their customer journey."
If your customer data isn't consolidated yet, and what you want from segmentation isn't a complex campaign list but the foundational layer that feeds your pricing and media decisions, your forecasting too, that's the fourth camp, and it's the case Kleene was built for. We've priced what assembling the alternative stack costs if you want to compare routes, and if you bring us your customer data, we'll tell you straight which camp you're actually shopping in.