Skip to main content

Product capabilities

Emotion intelligence, built for voice

Valence extracts frustration, urgency, calm, and hesitation from speech acoustics (prosody, pace, energy) without reading transcripts.

6 primary emotion classes
1.8s avg detection latency
F1 scored per emotion class

What Valence detects

Six primary emotion signals, each scored independently with a confidence value. Detection runs on speech acoustics, not transcript text.

Frustrated

Elevated pitch variance, clipped speech segments, rising terminal intonation. The signal most predictive of escalation and churn. Highest detection F1 score across our benchmark set.

Urgent

Compressed inter-syllable gaps, higher energy density in the first utterance of each turn. Common in billing disputes and outage-related contacts. Often precedes frustration if unaddressed.

Calm

Stable pitch range, consistent energy across utterances, lower speech rate. The baseline for a well-managed call. A shift from calm to frustrated in under two minutes is the primary escalation signature.

Hesitant

Increased pause frequency within utterances, raised pitch mid-phrase, lower confidence markers. Relevant for sales and renewals workflows where hesitation signals an objection before it is stated.

Satisfied

Falling terminal intonation, reduced speech energy in final utterances, longer turn closings. A positive post-resolution signal. Used in QA workflows to identify which handling styles end calls positively.

Neutral

No dominant affective signal detected. Represents a stable, factual call state. Neutral calls rarely require intervention. Useful as a baseline denominator for emotion trend reporting across teams.

Accuracy you can report on

Detection latency averages 1.8 seconds from the first spoken word, measured across an internal benchmark of 50,000 call segments from actual contact center recordings.

F1 scores vary by emotion class. Frustration is our strongest signal at F1 0.87 on the benchmark set. Hesitation and satisfaction are harder cases, particularly in low-audio-quality environments. We publish our per-class scores in the developer docs rather than quoting a single aggregate number.

Accuracy degrades gracefully in noisy conditions. We report a confidence value on every event. Your team can set a confidence threshold below which events are suppressed from the agent view.

Abstract visualization of signal confidence levels using color-filled geometric shapes

Three output shapes, one underlying model

Choose how you consume the emotion signal based on your team's workflow.

Real-time event stream

Continuous JSON events pushed via WebSocket or SSE. Each event includes emotion class, confidence score, valence, and arousal values. Latency from first audio frame to first event averages 1.8 seconds.

Call-level summary

A post-call JSON report covering the emotion timeline, peak frustration score, duration of each detected state, and a normalized sentiment trend curve. Delivered within 30 seconds of call end via webhook.

Webhook payload

A configurable threshold-based push notification. When caller frustration exceeds your configured score, Valence fires a POST to your endpoint. Use it to trigger supervisor alerts, CRM flags, or agent guidance prompts.

post-call-summary.json
{
  "call_id": "cid_8f2a91b3",
  "duration_s": 342,
  "peak_frustration": 0.81,
  "peak_frustration_at": 127,
  "dominant_emotion": "frustrated",
  "emotion_timeline": [
    { "t": 0, "e": "neutral", "c": 0.72 },
    { "t": 45, "e": "urgent", "c": 0.64 },
    { "t": 127, "e": "frustrated", "c": 0.81 }
  ]
}

Privacy built into the processing model

Audio is analyzed in-flight, not stored. We process acoustics, not content. Your call recordings stay yours.

Processed on the wire

Audio frames are analyzed as they arrive and discarded immediately after scoring. We never buffer a full recording. There is nothing to subpoena, breach, or accidentally retain past your configured retention window.

Acoustics, not transcripts

Valence detects emotion from how speech sounds, not what is said. No transcription occurs in our pipeline. PII in the conversation (names, account numbers, card details) never enters our system.

Your recordings stay yours

Valence runs alongside your existing recording infrastructure, not through it. You retain full custody of call recordings. Our output is metadata only: emotion events and confidence scores keyed to your call IDs.