phonesim
Overview
phonesim puts any audio through a simulated phone call. It is a local, seeded model of the signal path of a telephone call: VoIP transport, landline and mobile interconnects, the telephony codecs, packet loss, jitter, level control, background noise and clock drift. No telephony provider is involved, so you can test and train audio systems against phone calls without placing them.
Use it for anything that has to work over calls: speech recognition, speaker verification, watermark detection, speech enhancement.
The codecs are real. G.722, G.726, Opus and, with an AMR-capable ffmpeg build, AMR-NB and AMR-WB are encoded and decoded through ffmpeg or the libopus library; G.711 is the ITU-T segmented coder. Nothing is approximated: a profile whose codec your machine cannot run fails when it is built and says what to install.
Listen
A public-domain recording from LibriVox, before and after the default profile, voip_to_cellular_narrowband: a VoIP call delivered to a regular mobile phone.
Installation
phonesim needs Python 3.10 or newer. Install it in a virtual environment:
# Debian / Ubuntu (libavcodec-extra adds the AMR codecs the mobile profiles need)
sudo apt-get update
sudo apt-get install -y python3-venv ffmpeg libavcodec-extra libopus0
# macOS (Apple Silicon): brew install ffmpeg opus
python3 -m venv .venv
source .venv/bin/activate
pip install phonesim On Debian and Ubuntu this runs all seven profiles. libavcodec-extra adds the AMR codecs that the two mobile profiles, including the default voip_to_cellular_narrowband, and stress_multi_transcode need; it makes the ffmpeg build GPLv3. Homebrew's ffmpeg has no AMR, so on macOS the landline, internet and G.722 profiles run; the README shows how to get an AMR-capable ffmpeg there. phonesim info lists the profiles and which codecs your machine can run.
pip install phonesim pulls in PyTorch. On Linux the default PyTorch build includes CUDA: about 3 GB to download and 5.5 GB on disk. On a machine without an NVIDIA GPU, install the CPU build first (about 0.2 GB), then pip install phonesim:
pip install torch --index-url https://download.pytorch.org/whl/cpu On an Intel Mac, use Python 3.12 or older and run pip install "numpy<2" after installing phonesim: the last PyTorch release for Intel Macs needs NumPy 1.
Quick start
Python
No recording at hand? Download the original clip from the listening example and save it as input.wav.
from phonesim import PhoneCallSimulator, load_audio, save_audio
x, sr = load_audio("input.wav", sr=24000) # resampled to 24 kHz
sim = PhoneCallSimulator(profile="voip_opus_wideband")
y = sim(x, seed=1234) # reproducible degraded audio
save_audio("degraded.wav", y, sr=24000) This writes degraded.wav, 24 kHz mono. sim(x) expects 24 kHz input, which load_audio gives you; for audio at another rate, pass input_sample_rate= to PhoneCallSimulator. It accepts a NumPy array or a torch tensor shaped [T], [B, T] or [B, C, T] and returns the same type and shape at 24 kHz. With per_example=True, each row of a batch gets its own call, which is the mode for generating training data.
With an AMR-capable ffmpeg, profile="voip_to_cellular_narrowband" runs the same call path as the listening example above.
Command line
# the profiles and the codecs this machine can run
phonesim info
# one file through an internet call, then through a landline
phonesim run --in input.wav --out call.wav --profile voip_opus_wideband --seed 7
phonesim run --in input.wav --out pstn.wav --profile pstn_narrowband --seed 7
# a folder of clips
phonesim batch --in-dir clips/ --out-dir calls/ --profile voip_opus_wideband --seed 7Built-in call paths
Each profile models one call path.
| Profile | Call path | Codecs |
|---|---|---|
| pstn_narrowband | Landline: an analogue loop into a G.711 exchange | G.711 |
| pstn_g726 | Landline over G.726 ADPCM | G.726 |
| voip_opus_wideband | WebRTC / VoIP over the internet, average loss up to 5 % per call, in bursts | Opus |
| voip_g722_wideband | SIP HD voice | G.722 |
| voip_to_cellular_wideband | VoIP into a mobile HD voice call | Opus → AMR-WB |
| voip_to_cellular_narrowband | VoIP into a regular (non-HD) mobile call; the default | Opus → AMR-NB |
| stress_multi_transcode | Synthetic stress chain, not a real route | Opus → AMR-NB → G.711 → AMR-WB |
Swipe the table sideways to see all columns.
The six real call paths set the speech level and add background noise on the sending side, apply the band edges of their channel and end with clock drift. The four packet-based paths also lose frames, conceal them on the receiving side and adapt a playout buffer. print(PhoneCallSimulator(profile=...).describe()) shows the exact chain.
Reproducibility
Every call takes a seed. The same input, seed, profile version and codec build give identical output on one machine and torch thread count. Profiles are versioned: a bare name is the latest version, and name@1 pins one. sim(x, seed=1234, return_log=True) also returns a run log with the parameters drawn for each stage and the codec build that ran.
Limitations
- No EVS and no G.729. EVS has no open encoder and ffmpeg has no G.729 encoder, so VoLTE with EVS is not modelled.
- Transport effects are statistical. Packet loss, jitter-buffer adaptation and level control are parametric models, not RTP or WebRTC implementations.
- No acoustic path. Room reverberation, handset acoustics and echo are out of scope.
- No published comparison with recorded calls. Each profile is a physically motivated model of its path. The README describes how to check a profile against your own path.
