Retrieve results

Results at the file/call level

After the processing is complete, you can retrieve the results.

🔉 Audio files

🌐 cURL Example

curl --request GET \ 
     --url https://api.behavioralsignals.com/v5/detection/clients/<your-cid>/processes/<process-id>/results \
     --header 'X-Auth-Token: <your-api-token>' \
     --header 'accept: application/json'

🐍 Python SDK Example

Upload a file and wait for its results with:

from behavioralsignals import Client
from behavioralsignals.utils import print_results

client = Client()

process = client.deepfakes.upload_audio(file_path="audio.wav")
result = client.deepfakes.wait_for_result(
    pid=process.pid,
    timeout=600,
)

print_results(result.results)

Client() reads your credentials from the BEHAVIORALSIGNALS_CID and BEHAVIORALSIGNALS_API_KEY environment variables (see setup).

If the process has already finished, client.deepfakes.get_result(pid=...) returns the result without waiting.

result is an object of class ResultResponse, defined here

Response Example

{
  "pid": <your-pid>,
  "cid": <your-cid>,
  "code": 2,
  "message": "Processing Complete",
  "modelVersion": "7.7.0",
  "results": [
    {
      "id": "0",
      "startTime": "0.031",
      "endTime": "13.151",
      "task": "asr",
      "prediction": [
        {
          "label": " This is deep fake example of what is possible with powerful computer and editing. It took around seventy two hours to create this example from scratch using extremely powerful GPU. It could improve with more computing time, but ninety percent people cannot tell the difference"
        }
      ],
      "finalLabel": " This is deep fake example of what is possible with powerful computer and editing. It took around seventy two hours to create this example from scratch using extremely powerful GPU. It could improve with more computing time, but ninety percent people cannot tell the difference",
      "level": "utterance"
    },
    {
      "id": "0",
      "startTime": "0.031",
      "endTime": "13.151",
      "task": "diarization",
      "prediction": [
        {
          "label": "SPEAKER_00"
        }
      ],
      "finalLabel": "SPEAKER_00",
      "level": "utterance"
    },
    {
      "id": "0",
      "startTime": "0.031",
      "endTime": "13.151",
      "task": "language",
      "prediction": [
        {
          "label": "en",
          "posterior": "0.9990234375"
        },
        {
          "label": "es",
          "posterior": "7.838010787963867e-05"
        },
        {
          "label": "pt",
          "posterior": "7.361173629760742e-05"
        }
      ],
      "finalLabel": "en",
      "level": "utterance"
    },
    {
      "id": "0",
      "startTime": "0.031",
      "endTime": "13.151",
      "task": "deepfake",
      "prediction": [
        {
          "label": "spoofed",
          "posterior": "0.9997"
        },
        {
          "label": "bonafide",
          "posterior": "0.0003"
        }
      ],
      "finalLabel": "spoofed",
      "level": "utterance"
    }
  ]
}
  • The results object is a list of predictions - corresponding to each utterance detected in the audio file. The startTime and endTime mark the start and end of the utterance.
  • Each result item contains a set of tasks: asr, diarization, language and deepfake
  • The "deepfake" task consists of two classes: "bonafide" (authentic) and "spoofed" (deepfake).
  • The posterior denotes the probability of the corresponding class.
  • The finalLabel field always displays the dominant class.

🎥 Video files

🚧

Video deepfake detection is experimental

The endpoints and the response schema may still change, the models behind it are being tuned. Please get in touch before you build a production workflow on it.

A video process returns two separate results lists for the two modalities: audio_results and video_results.

To poll a video process before its results are ready, call https://api.behavioralsignals.com/v5/detection/clients/<your-cid>/processes/video/<process-id>, which returns the same status and statusmsg fields as an audio process.

🌐 cURL Example

curl --request GET \
     --url https://api.behavioralsignals.com/v5/detection/clients/<your-cid>/processes/video/<process-id>/results \
     --header 'X-Auth-Token: <your-api-token>' \
     --header 'accept: application/json'

🐍 Python SDK Example

Video support requires SDK version >= 0.5.0.

from behavioralsignals import Client
from behavioralsignals.utils import print_results

client = Client()

process = client.deepfakes.upload_video(file_path="video.mp4")
result = client.deepfakes.wait_for_video_result(
    pid=process.pid,
    timeout=600,
)

print_results(result.audio_results)
print_results(result.video_results)

If the process has already finished, client.deepfakes.get_video_result(pid=...) returns the result without waiting.

result is an object of class VideoResultResponse, defined here

Response Example

{
    "pid": <your-pid>,
    "cid": <your-cid>,
    "code": 2,
    "message": "Processing Complete",
    "modelVersion": "7.7.0",
    "audio_results": [
        {
            "id": "0",
            "startTime": "0.031",
            "endTime": "10.696",
            "task": "asr",
            "prediction": [
                {
                    "label": " It took around seventy-two hours to create this example from scratch using extremely powerful GPU. It could be improved with more computing time, but ninety percent of people cannot tell the difference."
                }
            ],
            "finalLabel": " It took around seventy-two hours to create this example from scratch using extremely powerful GPU. It could be improved with more computing time, but ninety percent of people cannot tell the difference.",
            "level": "utterance"
        },
        {
            "id": "0",
            "startTime": "0.031",
            "endTime": "10.696",
            "task": "diarization",
            "prediction": [
                {
                    "label": "SPEAKER_00"
                }
            ],
            "finalLabel": "SPEAKER_00",
            "level": "utterance"
        },
        {
            "id": "0",
            "startTime": "0.031",
            "endTime": "10.696",
            "task": "language",
            "prediction": [
                {
                    "label": "en",
                    "posterior": "0.99951171875"
                },
                {
                    "label": "fr",
                    "posterior": "6.812810897827148e-05"
                },
                {
                    "label": "es",
                    "posterior": "5.221366882324219e-05"
                }
            ],
            "finalLabel": "en",
            "level": "utterance"
        },
        {
            "id": "0",
            "startTime": "0.031",
            "endTime": "10.696",
            "task": "deepfake",
            "prediction": [
                {
                    "label": "spoofed",
                    "posterior": "0.9998"
                },
                {
                    "label": "bonafide",
                    "posterior": "0.0002"
                }
            ],
            "finalLabel": "spoofed",
            "level": "utterance"
        }
    ],
    "video_results": [
        {
            "id": "0",
            "startTime": "0.00",
            "endTime": "1.00",
            "task": "visual_deepfake",
            "prediction": [
                {
                    "label": "bonafide",
                    "posterior": "0.829"
                },
                {
                    "label": "spoofed",
                    "posterior": "0.171"
                }
            ],
            "finalLabel": "bonafide"
        },
        {
            "id": "1",
            "startTime": "1.00",
            "endTime": "2.00",
            "task": "visual_deepfake",
            "prediction": [
                {
                    "label": "bonafide",
                    "posterior": "0.8505"
                },
                {
                    "label": "spoofed",
                    "posterior": "0.1495"
                }
            ],
            "finalLabel": "bonafide"
        }, ...
    ]
}
  • The video_results object is a list of predictions corresponding to the video segments. The startTime and endTime mark the analyzed segment. Each result item contains the visual_deepfake task with the predicted posteriors and the finalLabel which can either be "bonafide" (authentic) or "spoofed" (deepfake).
  • The audio_results object has the same shape as the results of an audio process, with the asr, diarization, language and deepfake tasks per utterance. It is empty when the video has no audio track.
  • The picture and the speech are judged independently and they can disagree. A real recording with a cloned voice dubbed over it, or a face swap carrying authentic audio, is flagged on one side only.

Did this page help you?