Product capabilities
Emotion intelligence, built for voice
Valence extracts frustration, urgency, calm, and hesitation from speech acoustics (prosody, pace, energy) without reading transcripts.
What Valence detects
Six primary emotion signals, each scored independently with a confidence value. Detection runs on speech acoustics, not transcript text.
Frustrated
Elevated pitch variance, clipped speech segments, rising terminal intonation. The signal most predictive of escalation and churn. Highest detection F1 score across our benchmark set.
Urgent
Compressed inter-syllable gaps, higher energy density in the first utterance of each turn. Common in billing disputes and outage-related contacts. Often precedes frustration if unaddressed.
Calm
Stable pitch range, consistent energy across utterances, lower speech rate. The baseline for a well-managed call. A shift from calm to frustrated in under two minutes is the primary escalation signature.
Hesitant
Increased pause frequency within utterances, raised pitch mid-phrase, lower confidence markers. Relevant for sales and renewals workflows where hesitation signals an objection before it is stated.
Satisfied
Falling terminal intonation, reduced speech energy in final utterances, longer turn closings. A positive post-resolution signal. Used in QA workflows to identify which handling styles end calls positively.
Neutral
No dominant affective signal detected. Represents a stable, factual call state. Neutral calls rarely require intervention. Useful as a baseline denominator for emotion trend reporting across teams.
Accuracy you can report on
Detection latency averages 1.8 seconds from the first spoken word, measured across an internal benchmark of 50,000 call segments from actual contact center recordings.
F1 scores vary by emotion class. Frustration is our strongest signal at F1 0.87 on the benchmark set. Hesitation and satisfaction are harder cases, particularly in low-audio-quality environments. We publish our per-class scores in the developer docs rather than quoting a single aggregate number.
Accuracy degrades gracefully in noisy conditions. We report a confidence value on every event. Your team can set a confidence threshold below which events are suppressed from the agent view.
Three output shapes, one underlying model
Choose how you consume the emotion signal based on your team's workflow.
Real-time event stream
Continuous JSON events pushed via WebSocket or SSE. Each event includes emotion class, confidence score, valence, and arousal values. Latency from first audio frame to first event averages 1.8 seconds.
Call-level summary
A post-call JSON report covering the emotion timeline, peak frustration score, duration of each detected state, and a normalized sentiment trend curve. Delivered within 30 seconds of call end via webhook.
Webhook payload
A configurable threshold-based push notification. When caller frustration exceeds your configured score, Valence fires a POST to your endpoint. Use it to trigger supervisor alerts, CRM flags, or agent guidance prompts.
{
"call_id": "cid_8f2a91b3",
"duration_s": 342,
"peak_frustration": 0.81,
"peak_frustration_at": 127,
"dominant_emotion": "frustrated",
"emotion_timeline": [
{ "t": 0, "e": "neutral", "c": 0.72 },
{ "t": 45, "e": "urgent", "c": 0.64 },
{ "t": 127, "e": "frustrated", "c": 0.81 }
]
}
Privacy built into the processing model
Audio is analyzed in-flight, not stored. We process acoustics, not content. Your call recordings stay yours.
Processed on the wire
Audio frames are analyzed as they arrive and discarded immediately after scoring. We never buffer a full recording. There is nothing to subpoena, breach, or accidentally retain past your configured retention window.
Acoustics, not transcripts
Valence detects emotion from how speech sounds, not what is said. No transcription occurs in our pipeline. PII in the conversation (names, account numbers, card details) never enters our system.
Your recordings stay yours
Valence runs alongside your existing recording infrastructure, not through it. You retain full custody of call recordings. Our output is metadata only: emotion events and confidence scores keyed to your call IDs.