Definition of behaviors
Basic terminology of the behavioral outputs
Behavioral signals come from how something is said. For every utterance longer than 1 second, the Behavioral API returns 24 signals, grouped in 8 dimensions: 18 behavioral signals and 6 speaker signals (gender and age). It also returns a continuous emotion intensity score.
24 signals
emotion,strength,positivity,speaking_rate,hesitation,engagement,genderandageare dimensions that group the 24 signals. Each class label inside a dimension is a separate signal with its own probability.emotion: angry,engagement: withdrawnandhesitation: yesare three different findings.All eight dimensions are scored on every utterance longer than 1 second, at the same time. One utterance can be sad + weak + negative + slow + hesitating + withdrawn, with a voice estimated as female and aged 31 - 45. That gives 4 × 3 × 3 × 3 × 2 × 3 × 2 × 4 = 5,184 possible profiles per utterance.
📊 The 24 signals at a glance
| Dimension | What it measures | Signals | # |
|---|---|---|---|
emotion | The basic emotion in the voice | happy, angry, sad, neutral | 4 |
strength | Arousal: how much energy the voice carries | strong, weak, neutral | 3 |
positivity | Valence: the sentiment of the tone of voice | positive, negative, neutral | 3 |
speaking_rate | How fast the speaker talks, compared to speakers in general, not to their own pace | fast, slow, normal | 3 |
hesitation | Signs of hesitation in the speech | yes, no | 2 |
engagement | Whether the tone sounds involved or detached | engaged, withdrawn, neutral | 3 |
gender | The sex of the speaker, estimated from the voice | female, male | 2 |
age | The age range of the speaker, estimated from the voice | 18 - 22, 23 - 30, 31 - 45, 46 - 65 | 4 |
| Total | 24 |
Plus intensity: how intense the expressed emotion is, as a score between 0 and 1. It has no labels. It does not say which emotion is intense: for a score of anger, use the angry probability of emotion.
Active signals and baselines
Most signals point to something specific, such as emotion: angry, strength: strong or hesitation: yes. The rest are baselines: neutral (in emotion, strength, positivity and engagement), normal and no. Baselines matter too. A calm, neutral call is a finding in itself.
neutralmeans something different in each dimension
emotion: neutralmeans no clear emotion.strength: neutralmeans normal energy.positivity: neutralmeans neither positive nor negative.engagement: neutralmeans neither engaged nor withdrawn. Always read a label together with its dimension.
🗣️ The dimensions
Emotion
The basic emotion the speaker expresses through their voice.
- happy: joy, amusement, warmth or excitement.
- angry: irritation, frustration or anger.
- sad: sadness, disappointment or low mood.
- neutral: no clear emotion.
Strength
Arousal: how much energy and activation the voice carries, whatever the emotion. Anger and excitement are usually strong. Sadness and tiredness are usually weak.
- strong: high energy, an activated voice.
- weak: low energy, a flat or tired voice.
- neutral: normal energy.
Positivity
Valence: the sentiment of the speech, estimated from the tone of voice.
- positive: the tone is pleasant or favorable.
- negative: the tone is unpleasant or unfavorable.
- neutral: neither.
Positivity and emotion often agree, but not always. A calm complaint can be emotion: neutral and positivity: negative.
Speaking rate
How fast or slow the speaker talks, compared to speakers in general. The scale is global: the model learned it from a large speech corpus, and it does not adapt to each speaker's own pace.
- fast: faster than most speakers.
- slow: slower than most speakers.
- normal: a typical pace.
Hesitation
Whether the speech shows signs of hesitation, such as pauses, fillers ("uh", "um"), restarts or drawn-out words.
- yes: the speech shows signs of hesitation.
- no: no signs of hesitation.
Engagement
Whether the speaker sounds involved in the conversation.
- engaged: the speaker sounds interested and involved.
- withdrawn: the speaker sounds detached or uninterested.
- neutral: neither.
Gender
The sex of the speaker, estimated from the voice.
- female
- male
Age
The age range of the speaker, estimated from the voice.
- 18 - 22
- 23 - 30
- 31 - 45
- 46 - 65
Intensity
How intense the expressed emotion is, as a continuous score between 0 and 1. Higher means more intense. Unlike the dimensions above, intensity has no labels and no finalLabel. Its prediction holds a single score. It does not say which emotion is intense: for a score of anger, use the angry probability of emotion.
ℹ️ Reading a result
Each dimension comes back as its own result, one per utterance (or per segment, when streaming). The prediction list holds every signal of that dimension with its probability (posterior). The probabilities of one dimension add up to 1. finalLabel is the one answer the API gives for the dimension.
{
"id": "0",
"startTime": "0.031",
"endTime": "6.427",
"task": "engagement",
"prediction": [
{ "label": "withdrawn", "posterior": "0.7433" },
{ "label": "neutral", "posterior": "0.2427" },
{ "label": "engaged", "posterior": "0.014" }
],
"finalLabel": "withdrawn",
"level": "utterance"
}finalLabel is the label with the highest probability. Use finalLabel when you need one answer, and the probabilities when you need a score, for example to track how angry rises over a call.
All results of one utterance share its id, startTime and endTime. In batch results, gender and age are given per speaker: every utterance of a speaker longer than 1 second gets the same values, averaged over the file. When streaming, gender and age are estimated on each segment.
The probabilities are strings, so convert them to numbers first. Each one is a score from 0 to 1. positivity is valence and strength is arousal, the two dimensions often used in emotion research. Valence and arousal are often given on a scale from -1 to 1. To get that scale, subtract the probabilities of the two opposite labels: positive minus negative for valence, and strong minus weak for arousal. All probabilities are model estimates: use them to compare utterances and follow trends, not as exact measurements.
Short utterancesIn batch results, behavioral signals need utterances longer than 1 second. For shorter utterances, only
diarization,asrandlanguageare returned. When streaming, every segment gets all 24 signals andintensity, however short it is.
👥 Other outputs
Besides the 24 signals, the API also returns speech attributes with open-ended values: diarization (who speaks), asr (what is said) and language. Speaker IDs such as SPEAKER_00 are generic: the API does not say who is the agent and who is the customer. With embeddings enabled, it also returns speaker embeddings and behavioral embeddings (features). See Retrieve results for the full response schema.
🧑💻 Our Technology
Our models estimate each signal from how the voice sounds, not from the words. Each dimension has its own output, which is why every utterance gets a full profile and not a single label.
You can learn more on how to send a file in Submit audio file, or stream live audio with Streaming using the Python SDK.
Updated about 18 hours ago

