Discover / Video & YouTube

SpeechBrain

by speechbrainPython

All-in-one conversational AI and speech toolkit.

Repositorystable

Maturity: stable because 6y old, v1.1.0 released 126d ago. Derived from release and commit history, not a rating.

Stars
12k
Forks
1.7k
Downloads / mo
1.7M
Last commit
2026-06-15
License
Apache-2.0
Open issues
186

Market and trust evidence

Edition not yet matched

No exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.

Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.

In practice

Written by AI from this repository’s README · high confidence

Building speech recognition, speaker recognition, enhancement or separation systems means stitching separate codebases together.

Use it when

When you want one consistent toolkit for training or fine tuning speech models and running pretrained inference.

Not the right pick when

Wrong pick if you only need a single transcription binary, since it is a PyTorch research and training toolkit.

Capabilities

  • over 200 competitive training recipes on more than 40 datasets
  • training from scratch or fine tuning Whisper, Wav2Vec2, WavLM and Hubert
  • over 100 pretrained models hosted on HuggingFace
  • inference interfaces that transcribe a file in three lines of code
  • hyperparameters in YAML with training orchestrated by a Python script
  • over 30 tutorials covering the toolkit and Conversational AI

Cost: Free and open source

Install

Derived from the published package name in the repository, not from a model.

Video walkthroughs

Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.

What the repository ships

Has testsHas docsSecurity policyCI configured

Detected from the actual files in the repository root.

Latest release v1.1.0

Published 2026-03-30

This major release extends SpeechBrain's support for SpeechLLMs and introduces several new features, recipes, and improvements.

Highlights

  • Feature Caching — Save extracted features (e.g. wav2vec embeddings) to disk and load them on the fly, skipping recomputation. This powers our first ASR SpeechLLM recipe on LibriSpeech, enabling LLM-based training with pre-computed embeddings.
  • New Recipes — SpeechLLM for ASR and translation, streaming SSL, FocalCodec, and SENSE models.

Along with internal improvements and bug fixes. Here follows a changelog of the main changes (omitting some minor bugfixes):

What's Changed

  • Reforging LLaMA lobe (code from Samsung AI Center-Cambridge) by @TParcollet in https://github.com/speechbrain/speechbrain/pull/2850
  • Streaming recipe for BestRQ by @Chaanks in https://github.com/speechbrain/speechbrain/pull/2790
  • Librilight data preparation for SpeechBrain SSL (code from Samsung AI Center Cambridge) by @shucongzhang in https://github.com/speechbrain/speechbrain/pull/2765
  • Alignment with CTC ASR models powered by k2. by @ZhaoZeyu1995 in https://github.com/speechbrain/speechbrain/pull/2772
  • SpeechLLM (with LLaMA) and Conformer recipe for speech translation on CoVoST (Code from Samsung AI Center Cambridge) by @TParcollet in https://github.com/speechbrain/speechbrain/pull/2865
  • Feature caching proposal: CachedDynamicItem by @pplantinga in https://github.com/speechbrain/speechbrain/pull/2985
  • Adapting transducer greedy decoding by @younessdkhissi in https://github.com/speechbrain/speechbrain/pull/2975
  • Replace torchaudio I/O with soundfile-based audio_io wrapper by @Copilot in https://github.com/speechbrain/speechbrain/pull/2989
  • FocalCodec [NeurIPS 2025] by @lucadellalib in https://github.com/speechbrain/speechbrain/pull/3000
  • Caching: add compression + filename + closing/loading by @Adel-Moumen in https://github.com/speechbrain/speechbrain/pull/3005
  • Implement per-key padding configuration in PaddedBatch by @Adel-Moumen in https://github.com/speechbrain/speechbrain/pull/3008
  • SpeechLLM LibriSpeech recipe by @Adel-Moumen in https://github.com/speechbrain/speechbrain/pull/2885
  • Remove CTC CUDA + Move transducer loss in integrations by @Adel-Moumen in https://github.com/speechbrain/speechbrain/pull/3028
  • Adding SENSE models by @MaryemBouziane in https://github.com/speechbrain/speechbrain/pull/2998

New Contributors

  • @emmanuel-ferdman made their first contribution in https://github.com/speechbrain/speechbrain/pull/2900
  • @ofiryaish made their first contribution in https://github.com/speechbrain/speechbrain/pull/2871
  • @omidiu made their first contribution in https://github.com/speechbrain/speechbrain/pull/2923
  • @svecjan made their first contribution in https://github.com/speechbrain/speechbrain/pull/2934
  • @nouranali made their first contribution in https://github.com/speechbrain/speechbrain/pull/2855
  • @OscarFree made their first contribution in https://github.com/speechbrain/speechbrain/pull/2988
  • @younessdkhissi made their first contribution in https://github.com/speechbrain/speechbrain/pull/2975
  • @jordanozang made their first contribution in https://github.com/speechbrain/speechbrain/pull/2982
  • @jrochdi made their first contribution in https://github.com/speechbrain/speechbrain/pull/2947
  • @seohyunjun made their first contribution in https://github.com/speechbrain/speechbrain/pull/2996
  • @Daheer made their first contribution in https://github.com/speechbrain/speechbrain/pull/3022
  • @raotnameh made their first contribution in https://github.com/speechbrain/speechbrain/pull/2617
  • @Mr-Neutr0n made their first contribution in https://github.com/speechbrain/speechbrain/pull/3029
  • @georgesabr made their first contribution in https://github.com/speechbrain/speechbrain/pull/3038
  • @MaryemBouziane made their first contribution in https://github.com/speechbrain/speechbrain/pull/2998

Full Changelog: https://github.com/speechbrain/spe

Tags

README

<p align="center">

<img src="https://raw.githubusercontent.com/speechbrain/speechbrain/develop/docs/images/speechbrain-logo.svg" alt="SpeechBrain Logo"/>

</p>

Typing SVG

| 📘 Tutorials | 🌐 Website | 📚 Documentation | 🤝 Contributing | 🤗 HuggingFace | ▶️ YouTube | 🐦 X |

GitHub Repo stars Please, help our community project. Star on GitHub!

Exciting News (January, 2024): Discover what is new in SpeechBrain 1.0 here!

#

🗣️💬 What SpeechBrain Offers

  • SpeechBrain is an open-source PyTorch toolkit that accelerates Conversational AI development, i.e., the technology behind speech assistants, chatbots, and large language models.
  • It is crafted for fast and easy creation of advanced technologies for Speech and Text Processing.

🌐 Vision

  • With the rise of deep learning, once-distant domains like speech processing and NLP are now very close. A well-designed neural network and large datasets are all you need.
  • We think it is now time for a holistic toolkit that, mimicking the human brain, jointly supports diverse technologies for complex Conversational AI systems.
  • This spans speech recognition, speaker recognition, speech enhancement, speech separation, language modeling, dialogue, and beyond.
  • Aligned with our long-term goal of natural human-machine conversation, including for non-verbal individuals, we have recently added support for the EEG modality.

📚 Training Recipes

  • We share over 200 competitive training recipes on more than 40 datasets supporting 20 speech and text processing tasks (see below).
  • For any task, you train the model using these commands:

python train.py hparams/train.yaml
  • The hyperparameters are encapsulated in a YAML file, while the training process is orchestrated through a Python script.
  • We maintained a consistent code structure across different tasks.
  • For better replicability, training logs and checkpoints are hosted on Dropbox.

<a href="https://huggingface.co/speechbrain" target="_blank"> <img src="https://huggingface.co/front/assets/huggingface_logo.svg" alt="drawing" width="40"/> </a> Pretrained Models and Inference

  • Access over 100 pretrained models hosted on HuggingFace.
  • Each model comes with a user-friendly interface for seamless inference. For example, transcribing speech using a pretrained model requires just three lines of code:

from speechbrain.inference import EncoderDecoderASR

asr_model = EncoderDecoderASR.from_hparams(source="speechbrain/asr-conformer-transformerlm-librispeech", savedir="pretrained_models/asr-transformer-transformerlm-librispeech")
asr_model.transcribe_file("speechbrain/asr-conformer-transformerlm-librispeech/example.wav")

<a href="https://speechbrain.github.io/" target="_blank"> <img src="https://upload.wikimedia.org/wikipedia/commons/thumb/d/d0/Google_Colaboratory_SVG_Logo.svg/1200px-Google_Colaboratory_SVG_Logo.svg.png" alt="drawing" width="50"/> </a> Documentation

  • We are deeply dedicated to promoting inclusivity and education.
  • We have authored over 30 tutorials that not only describe how SpeechBrain works but also help users familiarize themselves with Conversational AI.
  • Every class or function has clear explanations and examples that you can run. Check out the documentation for more details 📚.

🎯 Use Cases

  • 🚀 Research Acceleration: Speeding up academic and industrial research. You can develop and integrate new models effortlessly, comparing their performance against our baselines.
  • ⚡️ Rapid Prototyping: Ideal for quick prototyping in time-sensitive projects.

#

🚀 Quick Start

To get started with SpeechBrain, follow these simple steps:

🛠️ Installation

Install via PyPI

  1. Install SpeechBrain using PyPI:

    pip install speechbrain
  1. Access SpeechBrain in your Python code:

    import speechbrain as sb

Install from GitHub

This installation is recommended for users who wish to conduct experiments and customize the toolkit according to their needs.

  1. Clone the GitHub repository and install the requirements:

    git clone https://github.com/speechbrain/speechbrain.git
    cd speechbrain
    pip install -r requirements.txt
    pip install --editable .
  1. Access SpeechBrain in your Python code:

    import speechbrain as sb

Any modifications made to the speechbrain package will be automatically reflected, thanks to the --editable flag.

✔️ Test Installation

Ensure your installation is correct by running the following commands:


pytest tests
pytest --doctest-modules speechbrain

🏃‍♂️ Running an Experiment

In SpeechBrain, you can train a model for any task using the following steps:


cd recipes/<dataset>/<task>/
python experiment.py params.yaml

The results will be saved in the output_folder specified in the YAML file.

📘 Learning SpeechBrain

  • Documentation: Detailed information on the SpeechBrain API, contribution guidelines, and code is available in the documentation.

#

🔧 Supported Technologies

  • SpeechBrain is a versatile framework designed for implementing a wide range of technologies within the field of Conversational AI.
  • It excels not only in individual task implementations but also in combining various technologies into complex pipelines.

🎙️ Speech/Audio Processing

| Tasks | Datasets | Technologies/Models |

| ------------- |-------------| -----|

| Speech Recognition | AISHELL-1, CommonVoice, DVoice, LibriSpeech, MEDIA, RescueSpeech, Switchboard, TIMIT, Tedlium2, Voicebank | CTC, Transducers, Transformers, Seq2Seq, Beamsearch techniques for CTC,seq2seq,transducers), Rescoring, Conformer, Branchformer, Hyperconformer, Kaldi2-FST |

| Speaker Recognition | VoxCeleb | ECAPA-TDNN, ResNET, Xvectors, PLDA, Score Normalization |

| Speech Separation | WSJ0Mix, LibriMix, WHAM!, WHAMR!, Aishell1Mix, BinauralWSJ0Mix | SepFormer, RESepFormer, SkiM, DualPath RNN, ConvTasNET |

| Speech Enhancement | DNS, Voicebank | SepFormer, MetricGAN, MetricGAN-U, SEGAN, spectral masking, time masking |

| Interpretability | ESC50 | Listenable Maps for Audio Classifiers (L-MAC), Learning-to-Interpret (L2I), Non-Negative Matrix Factorization (NMF), PIQ |

| Speech Generation | AudioMNIST | Diffusion, Latent Diffusion |

| Text-to-Speech | LJSpeech, LibriTTS | Tacotron2, Zero-Shot Multi-Speaker Tacotron2, FastSpeech2 |

| Vocoding | LJSpeech, LibriTTS | HiFiGAN, DiffWave

| Spoken Language Understanding | MEDIA, SLURP, Fluent Speech Commands, Timers-and-Such | Direct SLU, Decoupled SLU, Multistage SLU |

| Speech-to-Speech Translation | CVSS | Discrete Hubert, HiFiGAN, wav2vec2 |

| Speech Translation | Fisher CallHome (Spanish), IWSLT22(lowresource) | wav2vec2 |

| Emotion Classification | IEMOCAP, ZaionEmotionDataset | ECAPA-TDNN, wav2vec2, Emotion Diarization |

| Language Identification | VoxLingua107, CommonLanguage| ECAPA-TDNN |

| Voice Activity Detection | LibriParty | CRDNN |

| Sound Classification | ESC50, UrbanSound | CNN14, ECAPA-TDNN |

| Self-Supervised Learning | CommonVoice, LibriSpeech | wav2vec2 |

| Metric Learning | REAL-M, Voicebank | Blind SNR-Estimation, PESQ Learning |

| Alignment | TIMIT | CTC, Viterbi, Forward Forward |

| Diarization | AMI | ECAPA-TDNN, X-vectors, Spectral Clustering |

📝 Text Processing

| Tasks | Datasets | Technologies/Models |

| ------------- |-------------| -----|

| Language Modeling | CommonVoice, LibriSpeech| n-grams, RNNLM, [TransformerLM](https://arxiv.org/abs/170

Truncated. Read the full README on GitHub ↗

Related tools