#OpenAI-v3-large
OpenAI, Whisper 모델도 업그레이드 해주면 좋을듯. 팟플레이어에서 Faster-Whisper-XXL, large-v3 로 자막 종종 만들어서 보는데, 오류 좀 있더군요
September 2, 2026 at 12:37 PM
OpenAI, Whisper 모델도 업그레이드 해주면 좋을듯. 팟플레이어에서 Faster-Whisper-XXL, large-v3 로 자막 종종 만들어서 보는데, 오류 좀 있더군요
September 2, 2026 at 12:30 PM
Cohere zaprezentowało Transcribe – model ASR z 2 miliardami parametrów, który detronizuje Whisper Large V3 w szybkości i precyzji, radząc sobie z wyzwaniami dialektów arabskich.
Cohere rzuca wyzwanie OpenAI. Nowy model Transcribe Arabic i English wyznacza standardy
Cohere zaprezentowało Transcribe – model ASR z 2 miliardami parametrów, który detronizuje Whisper Large V3 w szybkości i precyzji, radząc sobie z wyzwaniami dialektów arabskich.
aisight.pl
July 7, 2026 at 6:05 PM
Oh we're doing this one again?

This is IDENTICAL to what OpenAI and Anthropic were saying about DeepSeek V3 and R1

This is an admission from Western AI that their 'Moat' is SO WEAK that you can simply 'steal' a large portion of that value by training on its exhaust.

THAT'S NOT GOOD FOR YOU DARIO!
June 25, 2026 at 2:55 PM
Snap Video Translator 1.1.1
Description: Snap Video Translator brings every step of localizing a video into a single Windows app. Load a video and it takes you straight through transcription, AI translation, and either subtitle burn-in or an AI voice-over. Transcription (runs locally)Runs OpenAI Whisper on your own PC. Choose from small / medium / large-v3-turbo / large-v3 to balance speed and accuracy. Your audio is never sent to an external server for transcription. AI translationWorks with AI services such as Gemini, OpenAI, and Claude, or with an OpenAI-compatible local LLM. Produces natural translations of your video's content. Subtitle burn-inAdjust font size, text color, outline, and position, then burn subtitles directly into the video. AI dubbingConverts the translation into speech and generates a voice-over in 7 languages (Japanese, English, Chinese, Korean, French, German, Spanish). Mix it with the original audio or replace it. Subtitle file exportExport subtitle files in SRT / VTT format. Review & edit modeCheck the subtitle text on screen before burn-in and edit it if needed. Batch processingProcess multiple videos at once. Who it's for- Video creators who want to reach viewers in other languages- Education and training teams localizing lectures and tutorials- Companies localizing internal videos Before you start- Transcription runs locally with Whisper; your audio is not sent anywhere.- Using a cloud AI service for translation requires that service's API key. You can also use an OpenAI-compatible local LLM.- AI translation and AI dubbing require an internet connection (no API key is needed for dubbing).- Subtitles and dubbed audio are AI-generated. Please review the output before public or business use.- FFmpeg is bundled as the video processing engine. System RequirementsWindows 10 version 19041.0 or higher (64-bit) * Japanese / English UI, selected automatically. Release Name: Snap Video Translator 1.1.1Size: 395.6 MBLinks: HOMEPAGE – NFO – Torrent Search Download: RAPiDGATOR
rlsbb.ru
June 6, 2026 at 9:36 PM
No idea, but heres the model card huggingface.co/openai/whisp...
openai/whisper-large-v3 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
May 24, 2026 at 11:03 PM
OpenRouter has literally spent time adopting OpenAI GPT-4o-mini Transcribe and Whisper Large V3 Turbo this (very young) May, as if there's nothing else to do.

As you can here, all of that makes me just a little bit salty.
May 1, 2026 at 11:11 PM
microsoft just dropped 3 MAI models: transcription, voice, and image. MAI-Transcribe-1 beats whisper-large-v3 on all 25 target languages.

the real story: microsoft is quietly building their way out of OpenAI dependency.
April 3, 2026 at 7:02 AM
Whisper Hugging Face Inference API bug
<p>Please, help!</p> <p>Although the default Whisper API setting is <strong>Transcribing</strong>, I receive a <strong>Translation</strong> (besides into different random languages!?!). Both whisper-large-v3 and whisper-large-v3-turbo. Why?</p> <p>final request = http.Request(‘POST’, url); <a href="//router.huggingface.co/hf-inference/models/openai/whisper-large-v3-turbo">//router.huggingface.co/hf-inference/models/openai/whisper-large-v3-turbo</a><br /> request.headers[‘Authorization’] = ‘Bearer $_hfToken’;<br /> request.headers[‘Content-Type’] = ‘audio/m4a’;<br /> request.bodyBytes = audioBytes;</p> <p>*** I tried to add additional headers:</p> <p>request.headers[‘Accept-Language’] = ‘en,en-US’;<br /> request.headers[‘language’] = ‘en’;<br /> request.headers[‘task’] = ‘transcribe’;</p> <p>… but it didn’t help.</p> <p>*** I also tried through json payload like:<br /> {<br /> ‘inputs’: base64Audio,<br /> ‘parameters’: {<br /> ‘task’: ‘transcribe’,<br /> ‘language’: ‘en’,<br /> },<br /> }</p> <p>… but returned an error "unexpected keyword argument ‘task’ "… yes ‘task’ is not in the public HF api’s parameters but it’s strange that there are no standard ones of the Whisper API.</p> <p>How can I fix this issue? I need to receive always and only Transcribe (not Translate to random language) as stated as the default behavior in Whisper Hugging Face Inference API.</p>
discuss.huggingface.co
February 5, 2026 at 10:27 PM
Whisper Large-v3 turns speech into text across 99 languages with scary-good accuracy - 6.7M downloads later, it's basically the Google Translate of audio transcription

🔗 https://huggingface.co/openai/whisper-large-v3

#SpeechRecognition #Multilingual #OpenAI
January 29, 2026 at 7:01 AM
Signal or noise? > AI, Vol. 6, Pages 297: Can Open-Source Large Language Models
Detect Medical Errors in Real-World Ophthalmology Reports? >> Comment below! #IoT #mhealth #industry40 #healthtech #AI
AI, Vol. 6, Pages 297: Can Open-Source Large Language Models Detect Medical Errors in Real-World Ophthalmology Reports?
Accurate documentation is critical in ophthalmology, yet clinical notes often contain subtle errors that can affect decision-making. This study prospectively compared contemporary large language models (LLMs) for detecting clinically salient errors in emergency ophthalmology encounter notes and generating actionable corrections. 129 de-identified notes, each seeded with a predefined target error, were independently audited by four LLMs (o3 (OpenAI, closed-source), DeepSeek-v3-r1 (Deepseek, open-source), MedGemma-27B (Google, open-source), and GPT-4o (OpenAI, closed-source)) using a standardized prompt. Two masked ophthalmologists graded error localization, relevance of additional issues, and overall recommendation quality, with within-case analyses applying appropriate nonparametric tests. Performance varied significantly across models (Cochran&rsquo;s Q = 71.13, p = 2.44 &times; 10&minus;15). o3 achieved the highest error localization accuracy at 95.7% (95% CI, 89.5&ndash;98.8), followed by DeepSeek-v3-r1 (90.3%), MedGemma-27b (80.9%), and GPT-4o (53.2%). Ordinal outcomes similarly favored o3 and DeepSeek-v3-r1 (both p &lt; 10&minus;9 vs. GPT-4o), with mean recommendation quality scores of 3.35, 3.05, 2.54, and 2.11, respectively. These findings demonstrate that LLMs can serve as accurate &ldquo;second-eyes&rdquo; for ophthalmology documentation. A proprietary model led on all metrics, while a strong open-source alternative approached its performance, offering potential for privacy-preserving on-premise deployment. Clinical translation will require oversight, workflow integration, and careful attention to ethical considerations.
dlvr.it
November 20, 2025 at 9:36 AM
The AI bubble was inflated because we live in a high-information, low-processing society, where we trust investors, the media and analysts to be far more critical than they are, leading to AI getting credit for what it "could do" rather than what it does today.
www.wheresyoured.at/premium-the-...
November 14, 2025 at 7:00 PM
New in JMIR MedEdu: Evaluating the Performance of DeepSeek-R1 and DeepSeek-V3 Versus OpenAI Models in the Chinese National Medical Licensing Examination: Cross-Sectional Comparative Study
Evaluating the Performance of DeepSeek-R1 and DeepSeek-V3 Versus OpenAI Models in the Chinese National Medical Licensing Examination: Cross-Sectional Comparative Study
Background: Deepseek-R1, an open-source large language model (LLM), has generated significant global interest in the past months. Objective: This study aimed to compare the performance of DeepSeek and OpenAI LLMs on the Chinese National Medical Licensing Examination (NMLE) and evaluate their potential in medical education #mededu. Methods: This cross-sectional study assessed 2 DeepSeek models (DeepSeek-R1 and DeepSeek-V3), 3 OpenAI models (ChatGPT-o1 pro, ChatGPT-o3 mini, and GPT-4o), and 2 additional Chinese LLMs (ERNIE 4.5 Turbo and Qwen 3) using the 2021 NMLE. Model performance was evaluated based on overall accuracy, accuracy across question types (A1, A2, A3 and A4, and B1), case analysis and non–case analysis questions, medical specialties, and accuracy consensus between different model combinations. Results: All LLMs successfully passed the NMLE. DeepSeek-R1 achieved the highest accuracy (573/597, 96%), followed by DeepSeek-V3 (558/600, 93%), both of which significantly outperformed ChatGPT-o1 pro (450/600, 75%), ChatGPT-o3 mini (455/600, 75.8%), and GPT-4o (452/600, 75.3%; P
dlvr.it
November 14, 2025 at 6:26 PM
So I'm running Hex, which is unfortunately Mac OS only, but it uses a local model. It uses a openai-whisper-large-v3-v20240903. In an ideal world, I would be running something similar on my iPhone as well, but I haven't found anything that runs a local model.

hex.kitlangton.com
HEX
VOICE → TEXT
hex.kitlangton.com
November 6, 2025 at 4:04 PM
specifically, it uses Whisper 3:
huggingface.co/openai/whisp...
I don't care if it's "useful" or a "less bad" use of AI, it's still AI and I'm not gonna support it. couldn't have used any of that 500k pounds raised on Kickstarter for localisation?
openai/whisper-large-v3 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
October 14, 2025 at 6:34 AM
実は進化している!ローカルで動くembeddingモデルたち

要約 日本語オンリーならruri v3 (わずか37mでOpenAIのtext-embedding-large-v3超え) もしかしたら日英だったらベターかも 多言語+コードならgranite-embedding はじめに LLMの普及からはや数年、175Bとかいう途方もないパラメータで動いていたLLMもいつの間にか4bに収まるようになり、スマホやPCで簡単に動かせるようになりました(現在だとQwen3-4b-thinking-2507などはかなり高性能です)。…
実は進化している!ローカルで動くembeddingモデルたち
要約 日本語オンリーならruri v3 (わずか37mでOpenAIのtext-embedding-large-v3超え) もしかしたら日英だったらベターかも 多言語+コードならgranite-embedding はじめに LLMの普及からはや数年、175Bとかいう途方もないパラメータで動いていたLLMもいつの間にか4bに収まるようになり、スマホやPCで簡単に動かせるようになりました(現在だとQwen3-4b-thinking-2507などはかなり高性能です)。 一方、embeddingモデルはといえば、OpenAIはtext-embeddings-3-small/la... Source link
inmobilexion.com
October 10, 2025 at 11:49 PM
今日のZennトレンド

実は進化している!ローカルで動くembeddingモデルたち
ローカル環境で動くオープンウェイトの埋め込みモデルが飛躍的に進化している。
日本語特化モデル「ruri-v3」はわずか37MのパラメータでOpenAIの大型モデルを超える高性能を発揮し、多言語・コード対応では「granite-embedding」が優秀である。
これらのモデルは、APIコスト削減や機密性保持を可能にし、ローカルでの高性能なRAG(検索拡張生成)システム構築に最適な選択肢を提供する。
実は進化している!ローカルで動くembeddingモデルたち
要約日本語オンリーならruri v3 (わずか37mでOpenAIのtext-embedding-large-v3超え)もしかしたら日英だったらベターかも多言語+コードならgranite-embedding はじめにLLMの普及からはや数年、175Bとかいう途方もないパラメータで動いていたLLMもいつの間にか4bに収まるようになり、スマホやPCで簡単に動かせるようになりました(現在だとQwen3
zenn.dev
October 10, 2025 at 9:18 PM