Definition of behaviors

Basic terminology of the behavioral outputs

Behavioral signals come from how something is said. For every utterance longer than 1 second, the Behavioral API returns 24 signals, grouped in 8 dimensions: 18 behavioral signals and 6 speaker signals (gender and age). It also returns a continuous emotion intensity score.

📘

24 signals

emotion, strength, positivity, speaking_rate, hesitation, engagement, gender and age are dimensions that group the 24 signals. Each class label inside a dimension is a separate signal with its own probability. emotion: angry, engagement: withdrawn and hesitation: yes are three different findings.

All eight dimensions are scored on every utterance longer than 1 second, at the same time. One utterance can be sad + weak + negative + slow + hesitating + withdrawn, with a voice estimated as female and aged 31 - 45. That gives 4 × 3 × 3 × 3 × 2 × 3 × 2 × 4 = 5,184 possible profiles per utterance.

📊 The 24 signals at a glance

DimensionWhat it measuresSignals#
emotionThe basic emotion in the voicehappy, angry, sad, neutral4
strengthArousal: how much energy the voice carriesstrong, weak, neutral3
positivityValence: the sentiment of the tone of voicepositive, negative, neutral3
speaking_rateHow fast the speaker talks, compared to speakers in general, not to their own pacefast, slow, normal3
hesitationSigns of hesitation in the speechyes, no2
engagementWhether the tone sounds involved or detachedengaged, withdrawn, neutral3
genderThe sex of the speaker, estimated from the voicefemale, male2
ageThe age range of the speaker, estimated from the voice18 - 22, 23 - 30, 31 - 45, 46 - 654
Total24

Plus intensity: how intense the expressed emotion is, as a score between 0 and 1. It has no labels. It does not say which emotion is intense: for a score of anger, use the angry probability of emotion.

Active signals and baselines

Most signals point to something specific, such as emotion: angry, strength: strong or hesitation: yes. The rest are baselines: neutral (in emotion, strength, positivity and engagement), normal and no. Baselines matter too. A calm, neutral call is a finding in itself.

⚠️

neutral means something different in each dimension

emotion: neutral means no clear emotion. strength: neutral means normal energy. positivity: neutral means neither positive nor negative. engagement: neutral means neither engaged nor withdrawn. Always read a label together with its dimension.

🗣️ The dimensions

Emotion

The basic emotion the speaker expresses through their voice.

  • happy: joy, amusement, warmth or excitement.
  • angry: irritation, frustration or anger.
  • sad: sadness, disappointment or low mood.
  • neutral: no clear emotion.

Strength

Arousal: how much energy and activation the voice carries, whatever the emotion. Anger and excitement are usually strong. Sadness and tiredness are usually weak.

  • strong: high energy, an activated voice.
  • weak: low energy, a flat or tired voice.
  • neutral: normal energy.

Positivity

Valence: the sentiment of the speech, estimated from the tone of voice.

  • positive: the tone is pleasant or favorable.
  • negative: the tone is unpleasant or unfavorable.
  • neutral: neither.

Positivity and emotion often agree, but not always. A calm complaint can be emotion: neutral and positivity: negative.

Speaking rate

How fast or slow the speaker talks, compared to speakers in general. The scale is global: the model learned it from a large speech corpus, and it does not adapt to each speaker's own pace.

  • fast: faster than most speakers.
  • slow: slower than most speakers.
  • normal: a typical pace.

Hesitation

Whether the speech shows signs of hesitation, such as pauses, fillers ("uh", "um"), restarts or drawn-out words.

  • yes: the speech shows signs of hesitation.
  • no: no signs of hesitation.

Engagement

Whether the speaker sounds involved in the conversation.

  • engaged: the speaker sounds interested and involved.
  • withdrawn: the speaker sounds detached or uninterested.
  • neutral: neither.

Gender

The sex of the speaker, estimated from the voice.

  • female
  • male

Age

The age range of the speaker, estimated from the voice.

  • 18 - 22
  • 23 - 30
  • 31 - 45
  • 46 - 65

Intensity

How intense the expressed emotion is, as a continuous score between 0 and 1. Higher means more intense. Unlike the dimensions above, intensity has no labels and no finalLabel. Its prediction holds a single score. It does not say which emotion is intense: for a score of anger, use the angry probability of emotion.

ℹ️ Reading a result

Each dimension comes back as its own result, one per utterance (or per segment, when streaming). The prediction list holds every signal of that dimension with its probability (posterior). The probabilities of one dimension add up to 1. finalLabel is the one answer the API gives for the dimension.

{
  "id": "0",
  "startTime": "0.031",
  "endTime": "6.427",
  "task": "engagement",
  "prediction": [
    { "label": "withdrawn", "posterior": "0.7433" },
    { "label": "neutral",   "posterior": "0.2427" },
    { "label": "engaged",   "posterior": "0.014" }
  ],
  "finalLabel": "withdrawn",
  "level": "utterance"
}

finalLabel is the label with the highest probability. Use finalLabel when you need one answer, and the probabilities when you need a score, for example to track how angry rises over a call.

All results of one utterance share its id, startTime and endTime. In batch results, gender and age are given per speaker: every utterance of a speaker longer than 1 second gets the same values, averaged over the file. When streaming, gender and age are estimated on each segment.

The probabilities are strings, so convert them to numbers first. Each one is a score from 0 to 1. positivity is valence and strength is arousal, the two dimensions often used in emotion research. Valence and arousal are often given on a scale from -1 to 1. To get that scale, subtract the probabilities of the two opposite labels: positive minus negative for valence, and strong minus weak for arousal. All probabilities are model estimates: use them to compare utterances and follow trends, not as exact measurements.

📘

Short utterances

In batch results, behavioral signals need utterances longer than 1 second. For shorter utterances, only diarization, asr and language are returned. When streaming, every segment gets all 24 signals and intensity, however short it is.

👥 Other outputs

Besides the 24 signals, the API also returns speech attributes with open-ended values: diarization (who speaks), asr (what is said) and language. Speaker IDs such as SPEAKER_00 are generic: the API does not say who is the agent and who is the customer. With embeddings enabled, it also returns speaker embeddings and behavioral embeddings (features). See Retrieve results for the full response schema.

🧑‍💻 Our Technology

Our models estimate each signal from how the voice sounds, not from the words. Each dimension has its own output, which is why every utterance gets a full profile and not a single label.

You can learn more on how to send a file in Submit audio file, or stream live audio with Streaming using the Python SDK.


Did this page help you?