1MiniMaxAI/MiniMax-H326,693 DOWNLOADS · 3,086 LIKESMiniMax published H3, a video generation model whose tags span text-to-video, image-to-video, video-to-video and joint audio-video generation. The same model is served through MiniMax's API platforms and its Hailuo web and desktop apps. The Hugging Face repository has collected 3,086 likes and about 26,700 downloads, and a cluster of community ComfyUI and LoRA conversions already sits behind it.MiniMax 发布 H3 视频生成模型,标签覆盖文生视频、图生视频、视频到视频,以及音视频联合生成。同一模型也通过 MiniMax 的 API 平台与海螺网页版、桌面端应用提供服务。该 Hugging Face 仓库已获 3,086 个点赞、约 26,700 次下载,社区围绕它的 ComfyUI 转换版与 LoRA 也已成群出现。MiniMax が動画生成モデル H3 を公開した。タグはテキストから動画、画像から動画、動画から動画、さらに音声付き動画の同時生成まで及ぶ。同じモデルは MiniMax の API プラットフォームや Hailuo のウェブ・デスクトップアプリでも提供されている。Hugging Face のリポジトリは 3,086 のいいねと約 26,700 ダウンロードを集め、その周辺にはコミュニティ製の ComfyUI 変換版や LoRA が並んでいる。
2inclusionAI/Ling-3.0-flash4,189 DOWNLOADS · 219 LIKESinclusionAI released Ling-3.0-flash, a native hybrid reasoning model with 124B total and 5.1B active parameters - roughly 12 percent of the size of the team's previous 1T-class flagship, which it says the new model matches or beats on key benchmarks. It uses a native hybrid-linear attention architecture and is listed on ModelScope and OpenRouter alongside Hugging Face.inclusionAI 发布 Ling-3.0-flash,一款原生混合推理模型,总参数 124B、激活参数 5.1B,约为该团队上一代万亿级旗舰的 12%,但官方称其在主要基准上持平甚至更好。模型采用原生混合线性注意力架构,除 Hugging Face 外也已上架 ModelScope 与 OpenRouter。inclusionAI がネイティブなハイブリッド推論モデル Ling-3.0-flash を公開した。総パラメータ 124B・アクティブ 5.1B と、同チームの前世代となる 1T 級フラッグシップの約12%の規模ながら、主要ベンチマークで同等以上だとしている。ネイティブなハイブリッド線形アテンション構成を採り、Hugging Face のほか ModelScope と OpenRouter でも提供されている。
3zai-org/GLM-5.22,480,368 DOWNLOADS · 4,902 LIKESZ.ai's GLM-5.2 is the lab's latest flagship for long-horizon tasks and the first in the line to deliver that capability across a full 1M-token context, a step up from GLM-5.1. It is by far the most used model in today's pool, with about 2.48 million downloads and 4,902 likes.智谱 Z.ai 的 GLM-5.2 是该实验室面向长时程任务的最新旗舰,也是该系列中首个在完整 100 万 token 上下文下提供这一能力的版本,相比 GLM-5.1 有明显提升。它是当日候选中使用量最高的模型,下载量约 248 万次,点赞 4,902 个。Z.ai の GLM-5.2 は、長期にわたるタスク向けの同ラボ最新フラッグシップであり、その能力を 100 万トークンのフルコンテキストで提供する初のモデルとして GLM-5.1 から前進した。本日の候補の中では突出して利用が多く、ダウンロードは約248万回、いいねは 4,902 に達している。
4lodestones/Kroma0 DOWNLOADS · 228 LIKESKroma v0.1 is a LoRA fine-tune of the Krea 2 image model, shipped as a single ComfyUI-compatible safetensors file of about 1.88 GB. It was rank-reduced to rank 256 and also carries the fully fine-tuned normalization and modulation tensors as weight deltas, so loading it on the base model reproduces the original fine-tune's behaviour.Kroma v0.1 是图像模型 Krea 2 的 LoRA 微调版本,以单个约 1.88 GB、兼容 ComfyUI 的 safetensors 文件发布。它被降秩至 rank 256,同时把完整微调后的归一化与调制张量作为权重增量一并携带,因此加载到基础模型上即可复现原始微调的表现。Kroma v0.1 は画像モデル Krea 2 の LoRA ファインチューンで、ComfyUI 対応の safetensors ファイル1つ(約1.88GB)として配布される。ランク256まで低ランク化したうえ、完全にファインチューンされた正規化・変調テンソルも重み差分として同梱しており、ベースモデルに読み込めば元のファインチューンの挙動を再現できる。
5SyzygyResearch/Mach-1-Additive-35B910 DOWNLOADS · 95 LIKESSyzygyResearch published Mach-1-Additive-35B, a ternary additive compression built on a Qwen3.5 mixture-of-experts base. Its card reports 95.0 percent mean score retention across 12 benchmarks for Mach-1 Small relative to the full-precision model, against 93.6 percent for a ternary Bonsai 27B and 85.6 percent for a Gemma 4 Q2_K_XL build.SyzygyResearch 发布 Mach-1-Additive-35B,基于 Qwen3.5 混合专家模型的三值「加性」压缩版本。模型卡显示,Mach-1 Small 在 12 项基准上相对全精度模型的平均得分保留率为 95.0%,而三值版 Bonsai 27B 为 93.6%,Gemma 4 Q2_K_XL 版本为 85.6%。SyzygyResearch が、Qwen3.5 の Mixture-of-Experts をベースにした三値の加算的圧縮モデル Mach-1-Additive-35B を公開した。モデルカードによれば、Mach-1 Small は12種のベンチマークで全精度モデルに対し平均95.0%のスコアを維持し、三値版 Bonsai 27B の93.6%、Gemma 4 Q2_K_XL の85.6%を上回るという。
2026-08-07 · FRIDAY · 13:06 PDT
1LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF332,992 DOWNLOADS · 423 LIKESA community finetune of Qwen3.6-35B-A3B, shipped in GGUF form for local inference, has pulled 332,992 downloads and 423 likes. The uploader describes it as an uncensored mixture-of-experts build with vision support. Its download count is far ahead of every other new entry on Hugging Face's trending list today.一个基于 Qwen3.6-35B-A3B 的社区微调模型以 GGUF 格式发布,面向本地推理,已获得 332,992 次下载和 423 个点赞。上传者称其为支持视觉输入、未经内容限制的混合专家模型。其下载量远超今日 Hugging Face 趋势榜上其他新条目。Qwen3.6-35B-A3Bをベースとしたコミュニティのファインチューン版が、ローカル推論向けのGGUF形式で公開され、332,992ダウンロードと423のいいねを集めた。投稿者は、画像入力に対応した無検閲のMoEモデルだと説明している。ダウンロード数は本日のHugging Faceトレンド一覧の他の新規エントリを大きく上回る。
2lightx2v/Minimax-h3-Turbo0 DOWNLOADS · 110 LIKESLightX2V published a distilled Turbo LoRA for MiniMax-H3, covering text-to-video, image-to-video and reference-to-video, with a GitHub repo for reproducing the results. Community ComfyUI ports already run it at four sampling steps. It has drawn 110 likes and a set of third-party conversions.LightX2V 发布了面向 MiniMax-H3 的蒸馏版 Turbo LoRA,覆盖文生视频、图生视频与参考图生视频,并提供了用于复现结果的 GitHub 仓库。社区的 ComfyUI 移植版已可在四步采样下运行。该模型已获得 110 个点赞,并出现了一批第三方转换版本。LightX2Vは、MiniMax-H3向けの蒸留版Turbo LoRAを公開した。テキストから動画、画像から動画、参照画像から動画までを対象とし、結果を再現するためのGitHubリポジトリも用意されている。コミュニティのComfyUI移植版はすでに4ステップのサンプリングで動作する。110のいいねを集め、第三者による変換版も複数登場している。
3Kijai/MiniMax-H3-TAE0 DOWNLOADS · 88 LIKESKijai trained a small 2D autoencoder for MiniMax-H3 so ComfyUI can show usable previews while a video generates. The author calls the result imperfect but better than latent2rgb, and now points users to a properly trained tiny video autoencoder from another developer. It runs through a preview-override node in ComfyUI-KJNodes.Kijai 为 MiniMax-H3 训练了一个小型二维自编码器,让 ComfyUI 在视频生成过程中能显示可用的预览。作者称效果并不理想,但优于 latent2rgb,并已在说明中指向另一位开发者训练的更完善的小型视频自编码器。它通过 ComfyUI-KJNodes 中的预览覆盖节点运行。Kijaiは、動画生成中のComfyUIで実用的なプレビューを表示できるよう、MiniMax-H3向けの小型2Dオートエンコーダを訓練した。出来は十分ではないがlatent2rgbよりは良いとしており、現在は別の開発者が適切に訓練した小型動画オートエンコーダを案内している。ComfyUI-KJNodesのプレビュー上書きノード経由で動作する。
4nvidia/Alpamayo2-Super1,591 DOWNLOADS · 88 LIKESNVIDIA released Alpamayo 2 Super, a 34B-parameter foundation model for autonomous-vehicle development that pairs a 32B vision-language backbone with a 2B diffusion expert. It is built to cover several AV development tasks in one model and forms part of the company's wider Alpamayo open platform.英伟达发布 Alpamayo 2 Super,这是一个面向自动驾驶开发的 340 亿参数基础模型,由 320 亿参数的视觉语言主干与 20 亿参数的扩散专家组合而成。它意在用单一模型覆盖多项自动驾驶开发任务,属于公司更大的 Alpamayo 开放平台的一部分。NVIDIAは、自動運転開発向けの340億パラメータ基盤モデルAlpamayo 2 Superを公開した。320億パラメータの視覚言語バックボーンと20億パラメータの拡散エキスパートを組み合わせた構成である。複数の自動運転開発タスクを一つのモデルで担うことを狙い、同社のAlpamayoオープンプラットフォームの一部を成す。
2026-08-06 · THURSDAY · 13:07 PDT
1larryvrh/MiniMax-H3-Turbo-Lora0 DOWNLOADS · 267 LIKESA community LoRA for the MiniMax-H3 video model that renders joint video plus synchronized stereo audio in 4 sampling steps instead of the usual roughly 20, about a 5x sampling speedup. The final checkpoint of this training round is sharp but carries known artifacts, and third-party ComfyUI conversions are already circulating.这是为 MiniMax-H3 视频模型制作的社区 LoRA,可在 4 个采样步内联合生成视频与同步立体声音频,而非通常的约 20 步,采样速度提升约 5 倍。本轮训练的最终检查点画面锐利但存在已知瑕疵,第三方的 ComfyUI 转换版本也已开始流传。MiniMax-H3ビデオモデル向けのコミュニティ製LoRAで、通常約20ステップのところを4サンプリングステップで映像と同期ステレオ音声を同時生成し、サンプリングを約5倍高速化する。今回の訓練ラウンドの最終チェックポイントは精細だが既知のアーティファクトがあり、サードパーティによるComfyUI変換版もすでに出回っている。
2LiquidAI/LFM2.5-2.6B-GGUF12,790 DOWNLOADS · 123 LIKESLiquidAI's official GGUF build of LFM2.5-2.6B, part of a new family of hybrid models designed for on-device deployment. LFM2.5 extends the LFM2 architecture with additional pre-training and reinforcement learning, and the GGUF packaging targets local inference with llama.cpp. The repo counts 12,790 downloads and 123 likes.LiquidAI 官方发布的 LFM2.5-2.6B GGUF 版本,属于面向端侧部署设计的新一代混合架构模型家族。LFM2.5 在 LFM2 架构基础上进行了扩展预训练与强化学习,GGUF 打包面向 llama.cpp 的本地推理。该仓库已有 12,790 次下载和 123 个赞。LiquidAI公式のLFM2.5-2.6B GGUF版で、オンデバイス展開向けに設計された新しいハイブリッドモデルファミリーの一つ。LFM2.5はLFM2アーキテクチャを追加事前学習と強化学習で拡張しており、GGUF形式はllama.cppでのローカル推論を想定する。リポジトリは12,790ダウンロードと123いいねを集めている。
3Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot272,963 DOWNLOADS · 106 LIKESA community-compiled collection of quantized and pruned MiniMax-H3 (Hailuo 3.0) weights in INT4, INT8, mixed, and NVFP4 formats for local inference environments such as ComfyUI. By unifying the formats in one structured repository it targets consumer GPUs with 16 to 24 GB of VRAM, and it has logged 272,963 downloads.这是社区整理的 MiniMax-H3(Hailuo 3.0)量化与剪枝权重合集,涵盖 INT4、INT8、混合精度与 NVFP4 格式,面向 ComfyUI 等本地推理环境。通过将多种格式统一到一个结构化仓库,它面向 16 至 24 GB 显存的消费级显卡,下载量已达 272,963 次。MiniMax-H3(Hailuo 3.0)の量子化・枝刈り済み重みをINT4、INT8、混合精度、NVFP4の各形式でまとめたコミュニティ編纂のコレクションで、ComfyUIなどのローカル推論環境向け。複数形式を一つの構造化リポジトリに統合することで16〜24GB VRAMのコンシューマーGPUを対象とし、ダウンロード数は272,963回に達している。
4sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP40 DOWNLOADS · 106 LIKESAn NVFP4 mixed-precision re-quantization of the Heretic uncensored Qwen3-VL-32B text encoder used with MiniMax-H3 video generation. At 15.7 GB the build fits on a single 16 GB card, continuing the push to squeeze the full H3 pipeline onto consumer hardware.这是 MiniMax-H3 视频生成所用 Heretic 无审查版 Qwen3-VL-32B 文本编码器的 NVFP4 混合精度再量化版本。文件体积 15.7 GB,可装入单张 16 GB 显卡,延续了把完整 H3 管线塞进消费级硬件的趋势。MiniMax-H3の動画生成で使われるHeretic(無検閲)版Qwen3-VL-32BテキストエンコーダーをNVFP4混合精度で再量子化したもの。15.7GBのビルドは16GBカード1枚に収まり、H3パイプライン全体をコンシューマーハードウェアに載せる流れを引き継いでいる。
5openai/whisper-large-v35,337,172 DOWNLOADS · 6,107 LIKESOpenAI's Whisper large-v3 is a state-of-the-art model for automatic speech recognition and speech translation, trained on more than 5 million hours of labeled data. Years after release it still ranks among the most-downloaded models on Hugging Face, with about 5.3 million downloads and over 6,100 likes.OpenAI 的 Whisper large-v3 是自动语音识别与语音翻译领域的先进模型,训练数据超过 500 万小时的标注音频。发布多年后,它仍位居 Hugging Face 下载量最高的模型之列,约有 530 万次下载和逾 6,100 个赞。OpenAIのWhisper large-v3は、500万時間超のラベル付きデータで学習された自動音声認識・音声翻訳の最先端モデル。リリースから年月を経た今もHugging Faceで最もダウンロードされるモデルの一つで、約530万ダウンロードと6,100超のいいねを集めている。
2026-08-05 · WEDNESDAY · 13:07 PDT
1ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot0 DOWNLOADS · 274 LIKESComfyUI conditioning encoders and generation tails built from an uncensored Heretic variant of Qwen3-VL-32B, shipped as BF16 and INT8 ConvRot builds. The package repurposes the vision-language model as H3 text and vision encoders for local generation workflows, and tops today's Hugging Face trending list at 274 likes.基于Qwen3-VL-32B无审查Heretic变体构建的ComfyUI条件编码器与生成尾部权重,提供BF16和INT8 ConvRot两种版本。该项目将这一视觉语言模型改造成H3文本与视觉编码器,用于本地生成工作流,以274个点赞位居今日Hugging Face趋势榜首。Qwen3-VL-32Bの無検閲Heretic派生モデルから構築されたComfyUI用条件付けエンコーダーと生成テールで、BF16とINT8 ConvRotの2形式で提供される。この視覚言語モデルをH3のテキスト・視覚エンコーダーとしてローカル生成ワークフローに転用したもので、274いいねで本日のHugging Faceトレンド首位に立つ。
2deepgrove/maple-preview0 DOWNLOADS · 136 LIKESDeepGrove's Maple-Preview is an open-source 20B-A1B ternary-weight reasoning model packed into a 5.31 GB checkpoint with 131K context. The team reports state-of-the-art reasoning for its weight class, IMO-level problem solving, and over 200 tokens per second on a Mac mini M4, five to sixteen times faster than efficient models like Gemma 4 and gpt-oss.DeepGrove的Maple-Preview是开源的20B-A1B三值权重推理模型,检查点仅5.31 GB,上下文长度13.1万token。团队称其推理能力在同量级中达到最先进水平,可解IMO级别的题目,在Mac mini M4上速度超过每秒200 token,比Gemma 4、gpt-oss等高效模型快5至16倍。DeepGroveのMaple-Previewは、5.31 GBのチェックポイントに収まる20B-A1Bの三値重み推論モデルで、13万1072トークンのコンテキストを持つオープンソースLLM。同クラス最高水準の推論性能とIMOレベルの問題解決力を掲げ、Mac mini M4で毎秒200トークン超と、Gemma 4やgpt-ossなど効率重視モデルの5〜16倍の速度をうたう。
3mistralai/Shieldstral-1.0-3B166 DOWNLOADS · 121 LIKESMistral's Shieldstral 1.0 is a compact 3B-parameter, policy-adaptive multimodal safety classifier. Rather than predicting fixed moderation categories, it scores text and image content against a safety policy written in natural language, letting guardrails be retargeted to new policies at inference time without retraining.Mistral的Shieldstral 1.0是一个紧凑的30亿参数、策略自适应多模态安全分类器。它不预测固定的审核类别,而是依据用自然语言书写的安全策略为文本和图像内容给出连续安全评分,使护栏无需重新训练即可在推理时切换到新策略。MistralのShieldstral 1.0は、30億パラメータのコンパクトなポリシー適応型マルチモーダル安全分類器。固定のモデレーションカテゴリを予測するのではなく、自然言語で記述された安全ポリシーに照らしてテキストや画像を連続スコアで評価し、再学習なしに推論時点で新しいポリシーへ切り替えられる。
4LGAI-EXAONE/K-EXAONE-2.0-750B-A37B325 DOWNLOADS · 129 LIKESLG AI Research released K-EXAONE 2.0, a frontier-scale multilingual mixture-of-experts model with 750B total and 37B active parameters. Scaled to more than three times its predecessor through upcycling, continual pretraining, and difficulty-focused mid-training, it is described as broadly competitive with leading open-weight models.LG AI研究院发布K-EXAONE 2.0,这是一个前沿规模的多语言混合专家模型,总参数7500亿,激活参数370亿。模型通过升级改造、持续预训练和以难度为核心的中期训练扩展到前代三倍以上规模,官方称其与领先的开放权重模型整体相当。LG AI Researchは、総パラメータ7500億、アクティブ370億のフロンティア級多言語Mixture-of-ExpertsモデルK-EXAONE 2.0を公開した。アップサイクリング、継続事前学習、難度重視の中間学習により前世代の3倍超の規模に拡張され、主要なオープンウェイトモデルと広く肩を並べる性能とされる。
5black-forest-labs/FLUX.1-dev546,530 DOWNLOADS · 13,996 LIKESBlack Forest Labs' FLUX.1-dev remains on Hugging Face's trending list long after release, with roughly 547,000 recent downloads and nearly 14,000 likes. The open-weight text-to-image model stays one of the most widely used bases for image generation and fine-tuning across the ecosystem.Black Forest Labs的FLUX.1-dev在发布许久之后仍停留在Hugging Face趋势榜上,近期下载约54.7万次,点赞近1.4万。这一开放权重文生图模型依然是社区中使用最广泛的图像生成与微调基础模型之一。Black Forest LabsのFLUX.1-devはリリースから時間が経った今もHugging Faceのトレンドに残り続けており、直近のダウンロードは約54万7000回、いいねは1万4000近くに上る。このオープンウェイトのテキスト画像生成モデルは、画像生成やファインチューニングの基盤として今も最も広く使われるモデルの一つだ。
2026-08-04 · TUESDAY · 13:04 PDT
1LiquidAI/LFM2.5-2.6B47,393 DOWNLOADS · 108 LIKESLFM2.5-2.6B is part of Liquid AI's LFM2.5 family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training, and the card claims tool use and instruction following competitive with models four times larger. The repo has passed 47,000 downloads.LFM2.5-2.6B 属于 Liquid AI 面向端侧部署设计的 LFM2.5 混合模型家族。它在 LFM2 架构基础上扩展到 128K 上下文窗口并加入智能体后训练,模型卡称其工具调用和指令遵循能力可与四倍规模的模型竞争。该仓库下载量已超过 4.7 万次。LFM2.5-2.6B は、Liquid AI がオンデバイス実行向けに設計したハイブリッドモデル群 LFM2.5 の一員。LFM2 アーキテクチャを基に 128K コンテキストウィンドウとエージェント向けポストトレーニングを加え、モデルカードではツール使用や指示追従で 4 倍規模のモデルに匹敵すると謳う。ダウンロード数は 4.7 万を超えた。
2realrebelai/MiniMax-H3_GGUFs40,010 DOWNLOADS · 97 LIKESA community GGUF packaging of MiniMax's H3 video generation model for ComfyUI, and the most downloaded of the H3 quant repos this week. It ships FL2VA and REF2V unet variants plus a quantized Qwen3-VL text encoder, pointing users to the official audio and video VAEs. Downloads have passed 40,000.这是社区为 ComfyUI 打包的 MiniMax H3 视频生成模型 GGUF 版本,也是本周下载量最高的 H3 量化仓库。它提供 FL2VA 和 REF2V 两类 unet 变体以及量化的 Qwen3-VL 文本编码器,并指引用户获取官方音频和视频 VAE。下载量已突破 4 万次。MiniMax の動画生成モデル H3 をコミュニティが ComfyUI 向けに GGUF 化したもので、H3 量子化リポジトリの中では今週最多ダウンロードを記録。FL2VA と REF2V の unet バリアントに加え量子化済み Qwen3-VL テキストエンコーダを同梱し、公式の音声・映像 VAE への導線も示す。ダウンロード数は 4 万を超えた。
3nvidia/NVIDIA-NemotronLabs-VoiceChat-11B80 DOWNLOADS · 77 LIKESNVIDIA NemotronLabs VoiceChat 11B is a speech model focused on natural conversational behavior. The card demonstrates smooth turn-taking with roughly 450 ms responses, barge-in handling where the model yields instantly when interrupted, and live tool calling, with code published in the NVIDIA NeMo Speech repository.NVIDIA NemotronLabs VoiceChat 11B 是一个专注自然对话行为的语音模型。模型卡演示了约 450 毫秒响应的流畅轮流对话、用户插话时模型即时让位的打断处理,以及实时工具调用,代码已发布在 NVIDIA NeMo Speech 仓库。NVIDIA NemotronLabs VoiceChat 11B は、自然な会話挙動に焦点を当てた音声モデル。モデルカードでは約 450 ミリ秒で応答する滑らかなターンテイキング、ユーザーの割り込みに即座に譲るバージイン処理、リアルタイムのツール呼び出しをデモしており、コードは NVIDIA NeMo Speech リポジトリで公開されている。
4huihui-ai/Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF14,046 DOWNLOADS · 76 LIKESAn uncensored build of DeepSeek-V4-Flash-0731 produced with abliteration to strip refusal behavior, distributed as GGUF files for llama.cpp. The maintainer describes it as a crude proof-of-concept for removing refusals without TransformerLens. It has drawn over 14,000 downloads.这是通过 abliteration 技术去除拒答行为的 DeepSeek-V4-Flash-0731 无审查版本,以 GGUF 格式发布,供 llama.cpp 使用。维护者称其为不依赖 TransformerLens 移除模型拒答的粗糙概念验证。下载量已超过 1.4 万次。アブリタレーションで拒否挙動を取り除いた DeepSeek-V4-Flash-0731 の無検閲版で、llama.cpp 向けの GGUF 形式で配布されている。メンテナは TransformerLens を使わずに拒否応答を除去する粗削りな概念実証と説明。ダウンロード数は 1.4 万を超えている。
5KRAFTON/A.X-K2-Raon-Speech-21B-A3B2,137 DOWNLOADS · 78 LIKESA.X K2 Raon-Speech is a bilingual English and Korean speech language model with about 21.2B total and 3.5B active parameters. Built on SK Telecom's A.X K2 Light mixture-of-experts text backbone, it adds KRAFTON's independently trained AuT speech encoder and a Mimi-style neural audio codec to unify speech understanding and generation in one model.A.X K2 Raon-Speech 是一个英韩双语语音语言模型,总参数约 212 亿,激活参数 35 亿。它以 SK 电讯的 A.X K2 Light 专家混合文本骨干为基础,加入 KRAFTON 独立训练的 AuT 语音编码器和 Mimi 风格神经音频编解码器,将语音理解与生成统一在单一模型中。A.X K2 Raon-Speech は英語と韓国語のバイリンガル音声言語モデルで、総パラメータ約 212 億、アクティブ 35 億。SK テレコムの MoE テキストバックボーン A.X K2 Light を土台に、KRAFTON が独自に学習した AuT 音声エンコーダと Mimi 方式のニューラル音声コーデックを組み合わせ、音声の理解と生成を単一モデルに統合している。
2026-08-03 · MONDAY · 13:05 PDT
1Comfy-Org/MiniMax-H32 DOWNLOADS · 420 LIKESComfy-Org's repackaging of MiniMaxAI's newly open-weight MiniMax-H3 model for ComfyUI. The repo ships the fl2va and ref2va diffusion checkpoints in bf16, int8, and fp8 variants, laid out in ComfyUI's expected folder structure for local generation workflows. It has drawn 420 likes since release.Comfy-Org 将 MiniMaxAI 新近开放权重的 MiniMax-H3 模型重新打包为 ComfyUI 可用格式。仓库提供 fl2va 和 ref2va 两类扩散检查点的 bf16、int8 和 fp8 版本,并按 ComfyUI 约定的目录结构组织,便于本地生成工作流使用。发布以来已获得 420 个点赞。Comfy-Org が、MiniMaxAI が新たにオープンウェイトで公開した MiniMax-H3 モデルを ComfyUI 向けに再パッケージしたもの。リポジトリには fl2va と ref2va の拡散チェックポイントが bf16、int8、fp8 の各バリアントで収録され、ローカル生成ワークフロー向けに ComfyUI 標準のフォルダ構成で配置されている。公開以来 420 のいいねを集めている。
2ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot0 DOWNLOADS · 80 LIKESAn INT8 ComfyUI repackaging of a community 'heretic' uncensored variant of Qwen3-VL-32B, built for use with MiniMax-H3 pipelines. It splits the model into a text and vision conditioning encoder covering language layers 0 to 49, plus an optional generation-only tail with the remaining layers and LM head for prompt enhancement.一个将社区'heretic'去审查版 Qwen3-VL-32B 以 INT8 格式重新打包供 ComfyUI 使用的仓库,面向 MiniMax-H3 流水线。它将模型拆分为覆盖语言层 0 至 49 的文本与视觉条件编码器,以及一个可选的仅生成尾部,后者包含其余层和语言模型头,用于提示词增强。コミュニティ製の「heretic」無検閲版 Qwen3-VL-32B を、MiniMax-H3 パイプラインでの利用に向けて INT8 形式で ComfyUI 用に再パッケージしたもの。モデルは言語層 0〜49 を含むテキスト・視覚条件付けエンコーダーと、残りの層と LM ヘッドを収めたプロンプト拡張用のオプションの生成専用テールに分割されている。
3unsloth/Inkling-Small-GGUF30,594 DOWNLOADS · 73 LIKESUnsloth's GGUF quantizations of Inkling-Small, a general-purpose multimodal mixture-of-experts model that accepts text, image, and audio input. The dynamic quants run down to a 1-bit UD-IQ1_S variant, and the repo has logged over 30,000 downloads, making it by far the most downloaded of the week's new model repos.Unsloth 为 Inkling-Small 制作的 GGUF 量化版本。该模型是一个通用多模态专家混合模型,支持文本、图像和音频输入。动态量化最低可至 1 比特的 UD-IQ1_S 版本,仓库下载量已超过 3 万次,是本周新模型仓库中下载量遥遥领先的一个。Unsloth による Inkling-Small の GGUF 量子化版。同モデルはテキスト・画像・音声入力に対応する汎用マルチモーダル MoE(専門家混合)モデルである。動的量子化は 1 ビットの UD-IQ1_S バリアントまで用意されており、リポジトリのダウンロード数はすでに 3 万回を超え、今週の新規モデルリポジトリの中で群を抜いて多い。
2026-08-02 · SUNDAY · 13:02 PDT
1nyralabs/CrisperWhisper2.0_large3,100 DOWNLOADS · 80 LIKESNyra Labs published CrisperWhisper 2.0 large, a Whisper-based model for verbatim speech recognition, on Hugging Face. It targets production use with multilingual support, word-level timestamps and explicit handling of disfluencies, plus CTranslate2 support. The model counts 3,100 downloads and 80 likes.Nyra Labs 在 Hugging Face 上发布了 CrisperWhisper 2.0 large,这是一个基于 Whisper 的逐字语音识别模型。它面向生产环境,支持多语言、词级时间戳以及对言语不流畅现象的显式处理,并支持 CTranslate2。该模型目前有 3100 次下载和 80 个点赞。Nyra Labs は、Whisper ベースの逐語的音声認識モデル CrisperWhisper 2.0 large を Hugging Face で公開した。多言語対応、単語単位のタイムスタンプ、言い淀みの明示的な処理を備えたプロダクション用途向けで、CTranslate2 にも対応する。ダウンロード数は 3,100 件、いいねは 80 件となっている。