One of the first questions we get from operations architects and engineering leads evaluating Valence is: what does integration actually look like? The concern is legitimate. Contact center telephony stacks are not simple environments. They involve multiple vendors, legacy SIP infrastructure, compliance requirements, and limited internal engineering bandwidth for integration projects. A tool that requires a six-month deployment is not a viable pilot candidate for most teams.
The answer depends on what CCaaS platform you are running and how your audio is routed. There are three integration paths, and each has a different profile in terms of setup time, latency, and operational complexity. Here is how to think about which one fits your situation.
Path 1: SIP Media Tap
A SIP media tap intercepts the audio stream at the SIP signaling layer, before it reaches the agent's soft phone, and routes a copy of the raw audio to the Valence inference endpoint. The original call path is unchanged. The tap is a passive fork, not an in-path component, which means adding it does not affect call quality or introduce single-point-of-failure risk into the live call routing.
This path gives you the lowest latency of the three options. Because you are working with raw audio arriving in near-real-time at the network layer, the audio delivery lag to the inference service is minimal. If your goal is sub-2-second frustration detection, a SIP tap is typically where you land.
The setup requirements: you need access to your SIP trunk configuration to add the media forking rule, and you need a network path from the media tap destination to the Valence webhook receiver. In most on-premises or hybrid SIP deployments, this is a configuration change that can be done in a few hours by someone with access to the SIP trunk settings. The harder constraint is sometimes organizational: getting change control approval for a SIP configuration modification can take longer than the technical work.
Where this path does not work: fully cloud-native CCaaS platforms that route audio through proprietary media servers with no SIP trunk exposure do not support media tapping at the network layer. You cannot tap what you cannot reach. If your CCaaS platform abstracts the telephony layer entirely, you are on Path 2 or Path 3.
Path 2: CCaaS Webhook Connector
Most modern CCaaS platforms offer some form of real-time event streaming or webhook-based audio delivery. The implementation varies significantly by platform. Some deliver audio segments as binary payloads in a streaming webhook. Others use a WebSocket connection that the external service subscribes to. A few expose a dedicated "call media streaming" API with its own SDK.
The integration model here is: Valence runs a receiver endpoint that the CCaaS platform is configured to send audio to as the call progresses. The CCaaS platform handles the audio segmentation and delivery timing. Valence processes each segment and emits an emotion event back through a configured outbound webhook to your CRM, supervisor dashboard, or internal event bus.
Setup effort for this path is typically lower than a SIP tap from a change control perspective, because it is a platform configuration within the CCaaS admin console rather than a network infrastructure change. Most CCaaS platform admins are comfortable with webhook configuration. The tradeoff is that you are bound by the CCaaS platform's audio delivery behavior, including its segment accumulation window and any throttling on the media streaming API.
The latency characteristics of Path 2 depend entirely on what the CCaaS platform does. Some platforms deliver audio in 500-millisecond chunks with minimal buffering. Others batch audio in 3-second windows before sending. If the platform you are running defaults to 3-second windows, that is your latency floor before Valence receives anything to analyze. Understanding this number is a critical part of evaluating whether a CCaaS webhook integration can meet your real-time requirements. Ask the platform vendor specifically about their media streaming segment window before assuming the integration will support live detection.
Path 3: Direct Audio Stream API
The third path is used when neither a SIP tap nor a CCaaS webhook integration is feasible, or when the contact center has already built custom call handling infrastructure and wants to embed emotion detection directly into that pipeline. In this model, your application sends raw audio to the Valence stream API directly, and receives emotion events on the same connection as a response stream.
This path has the most flexibility and the highest integration effort. It is the right choice for teams that are already running custom audio processing pipelines, building a conversational AI application that needs to embed emotion awareness, or deploying in an environment with non-standard telephony infrastructure. It is the wrong choice for a team that just wants to add emotion signals to an existing CCaaS deployment without writing code.
The stream API accepts PCM audio over a WebSocket connection, with configurable sample rate and channel layout. Valence returns a stream of emotion event objects in near-real-time as audio is processed. The event schema is consistent across all three integration paths, so if you start with Path 2 for a quick pilot and later want to move to a tighter Path 3 integration, the downstream event processing does not need to change.
Choosing the Right Path for Your Situation
The decision tree is simpler than it looks:
If you are running a SIP-based on-premises or hybrid platform and have access to your SIP trunk configuration, start with Path 1. It gives you the cleanest latency profile and the fewest dependencies on platform-specific APIs. The setup is a one-time configuration change with no ongoing maintenance.
If you are on a cloud CCaaS platform that offers media streaming webhooks, use Path 2. Check the segment window before committing to a latency target, and plan your pilot evaluation criteria around the platform's actual delivery behavior rather than theoretical minimums.
If you are building a custom application or need to integrate emotion detection at the code level, Path 3 is the right architectural choice. Budget for a proper integration sprint, not a configuration exercise.
What "No Six-Month Project" Actually Means
For Path 1 and Path 2 deployments on standard CCaaS platforms, the technical integration itself should take days, not months. The full pilot setup, including Valence account provisioning, integration configuration, and connecting the event output to a supervisor dashboard, is typically complete within a week for a small pilot team.
What can extend timelines is not the technical work. It is organizational process: change control for the SIP configuration, security review of the webhook endpoint, IT sign-off on the data path. These are real constraints and we work through them with every deployment partner. The point is that the technical complexity is not the limiting factor. If you are being told an integration is going to take six months, the question worth asking is whether that estimate is driven by the technical work or by the approval process layered around it.
A pilot scope should be genuinely small: one team, one queue, live emotion monitoring for a few weeks to validate detection quality on your specific call types and audio delivery path. Piloting at this scope avoids the organizational friction of a full deployment while generating the real-world data you need to make an informed decision about production rollout.
A Note on Audio Quality and Codec Constraints
One constraint that affects all three integration paths, and that is worth flagging before you start: the quality of the audio that reaches the inference service matters. Emotion detection relies on frequency-domain features, specifically prosodic and spectral characteristics, that degrade significantly under aggressive audio compression.
The most common problem we see is G.729 codec encoding on the telephony path, particularly in older SIP deployments. G.729 uses a 8 kbps narrow-band codec that discards frequency information above 4 kHz. This is sufficient for speech intelligibility, but it attenuates or eliminates some of the spectral features that carry reliable emotion signal. Detection accuracy on G.729 audio is measurably lower than on G.711 or wideband audio.
If you are running G.729, the question is whether you have the option to switch the codec for the audio path going to the Valence tap. You do not need wideband audio everywhere. You only need it on the copy of the audio that is going to the inference service. This is often a per-trunk or per-route configuration change rather than a full codec migration. Worth checking before assuming you are stuck with the constraint.