xClean Tools

AI Models — Daily Top 5

UPDATED 2026-09-06 01:03 PDT

2026-08-01 · SATURDAY · 13:03 PDT

  1. 1 unsloth/DeepSeek-V4-Flash-0731-GGUF 4,048 DOWNLOADS · 268 LIKES Unsloth published GGUF quantizations of DeepSeek's new V4-Flash-0731 checkpoint for local inference, drawing 4,048 downloads and 268 likes. The repo uses Unsloth's Dynamic 2.0 method with imatrix variants up to a 162GB Q8, and the model can also run in Unsloth Studio with toggles for High and Max thinking.
  2. 2 skt/A.X-K2 1,218 DOWNLOADS · 74 LIKES SK Telecom released A.X K2, a Mixture-of-Experts language model trained from scratch as an agentic foundation model and successor to A.X K1. It carries 688 billion total parameters with 33 billion active, uses a Think-Fusion training recipe, and supports English, Korean, and Chinese.

2026-07-31 · FRIDAY · 13:05 PDT

  1. 1 deepseek-ai/DeepSeek-V4-Flash-0731 0 DOWNLOADS · 869 LIKES DeepSeek shipped DeepSeek-V4-Flash-0731, the official release of V4-Flash that supersedes the preview version with substantially enhanced agentic capabilities. It keeps the same model structure, including an attached speculative decoding module, and outperforms DeepSeek-V4-Pro (Preview) on the benchmarks in its report. The V4-Flash line is a 284B-parameter MoE with 13B active and a one-million-token context.
  2. 2 thinkingmachines/Inkling-Small 2,971 DOWNLOADS · 186 LIKES Thinking Machines' Inkling-Small is a general-purpose multimodal MoE model that accepts text, image, and audio inputs and generates text, shipping in BF16 and NVFP4 with Tinker Cookbook support. It targets developers building agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation applications across multiple languages.
  3. 3 XYZAILab/XYZ-Aquila-mini 579 DOWNLOADS · 350 LIKES XYZ AI Lab's Aquila-mini is an open-weight thinking model for Deep Search, post-trained from Qwen3.6-35B-A3B through a bounded-exploration AI4AI pipeline. Humans define the target capability, constraints, risk boundaries, and acceptance policy, while AI agents diagnose failures and propose scoped changes.
  4. 4 XYZAILab/XYZ-Aquila-pro 869 DOWNLOADS · 326 LIKES Aquila-pro is the larger member of XYZ AI Lab's family of open-weight Deep Search agents, post-trained from Qwen3.5-397B-A17B through the same bounded-exploration AI4AI pipeline as its mini sibling. The recipe keeps humans in charge of capability targets and risk boundaries while AI agents drive iteration.
  5. 5 empero-ai/Qwythos-27B-v1 864 DOWNLOADS · 78 LIKES Empero's Qwythos-27B-v1 is an open-weight, full-parameter multimodal reasoning model, the larger sibling of Qwythos-9B trained on the same curriculum atop a Qwen3.5-27B base. It ships as a complete pre-RL checkpoint post-trained through SFT, DPO, and ESFT, with the stated goal that nothing was ablated to make it fit.

2026-07-30 · THURSDAY · 12:36 PDT

  1. 1 unsloth/Kimi-K3-GGUF 12,178 DOWNLOADS · 203 LIKES GGUF quantisations of Kimi K3, packaged so the model can be run locally through llama.cpp-based tooling rather than a hosted API. The repo covers a range from full-precision down to aggressive 4-bit builds, and the weights carry vision support.
  2. 2 Audio8/Audio8-TTS-Preview-0.6b 225 DOWNLOADS · 116 LIKES A 0.6-billion-parameter multilingual text-to-speech model with zero-shot voice cloning, released as a preview. The pitch is competitive speech quality at a size small enough to run without dedicated inference hardware.
  3. 3 Comfy-Org/Mage-Flow 44,714 DOWNLOADS · 95 LIKES Microsoft's Mage-Flow repackaged into single-file weights laid out for ComfyUI, so the model drops into an existing node graph without manual conversion. Downloads here run well ahead of the other entries on the list.
  4. 4 EschaLabs/Qwen3.6-35B-A3B-Escha-W2 201 DOWNLOADS · 85 LIKES A 2-bit quantised build of the Qwen3.6-35B-A3B mixture-of-experts model, produced with Escha Labs' own quantisation method. Pushing a 35B MoE to 2 bits is aimed at fitting it on hardware that could not otherwise hold it.
  5. 5 LiquidAI/LFM2.5-Encoder-230M 7,353 DOWNLOADS · 69 LIKES A 230-million-parameter multilingual bidirectional encoder built on the LFM2 architecture, released as part of a two-size family. Encoders of this kind serve retrieval and classification work rather than generation.

2026-07-29 · WEDNESDAY · 12:40 PDT

  1. 1 unsloth/Kimi-K3 410 DOWNLOADS · 163 LIKES Unsloth's quantized re-release of Moonshot AI's Kimi K3, the 2.8T parameter open-weight multimodal agentic model. The accompanying GGUF build points at a llama.cpp fork for running it, and notes that lossless Q8 weights come to 1.56TB.
  2. 2 nota-ai/Solar-Open2-250B-Nota-NVFP4 6,189 DOWNLOADS · 137 LIKES Nota AI's 4-bit quantization of Upstage's Solar Open2 250B, produced with a technique aimed at mixture-of-experts models. The weights use NVFP4 with a group size of 16, packed in compressed-tensors format for direct serving in vLLM.
  3. 3 microsoft/Mage-VL 702 DOWNLOADS · 91 LIKES Microsoft published a codec-native streaming multimodal model for image and video understanding, with a visual encoder trained from scratch at a compact 4B scale. The card frames it against a Moravec's paradox for vision-language models, which reason well offline yet stay slow and compute-heavy on simple real-time perception.
  4. 4 amd/Instella-MoE-16B-A3B-Think 730 DOWNLOADS · 77 LIKES AMD's fully open mixture-of-experts language model with 16 billion total and 2.8 billion active parameters, trained end to end from pre-training through reinforcement learning. It was trained from scratch on AMD Instinct MI300X and MI325X GPUs using the company's Primus framework.
  5. 5 LiquidAI/LFM2.5-Encoder-350M 5,327 DOWNLOADS · 66 LIKES Liquid AI released multilingual bidirectional encoders on the LFM2 architecture in two sizes, of which this 350M model is the larger one aimed at downstream quality. The 230M sibling targets tight latency and memory budgets instead.

2026-07-28 · TUESDAY · 12:40 PDT

  1. 1 moonshotai/Kimi-K3 99,214 DOWNLOADS · 7,840 LIKES Moonshot AI's Kimi K3 is an open-weight native multimodal agentic model of 2.8T parameters built on Kimi Delta Attention and Attention Residuals, with native vision and a 1-million-token context window. Downloads jumped from 2,850 to 99,214 over the past twelve hours as likes reached 7,840.
  2. 2 LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF 99,660 DOWNLOADS · 193 LIKES The sixth Genesis release of an uncensored, vision-capable fine-tune of the Qwen3.6-35B-A3B mixture-of-experts model, distributed as GGUF. The card frames its method as removing accumulated tensor noise from training. At 99,660 downloads it is the most pulled model on today's trending list.
  3. 3 microsoft/VibeVoice-ASR-BitNet 1,754 DOWNLOADS · 86 LIKES Microsoft's BitNet-quantized speech recognition model from the VibeVoice line, shipped in both safetensors and GGUF for the VibeASR.cpp runtime. The card links a technical report and an MIT license, and downloads have grown to 1,754.
  4. 4 ProCreations/grug-27b 1,170 DOWNLOADS · 70 LIKES A 27B Qwen3.5-based fine-tune trained to reason in terse, telegraphic notes instead of the usual verbose chain of thought. The card presents token efficiency as the point, contrasting a long deliberative trace with a clipped one that reaches the same decision.
  5. 5 microsoft/Mage-Flow-Turbo 1,582 DOWNLOADS · 71 LIKES A 4B-scale generative stack for native-resolution text-to-image work and instruction-based editing. The card argues that co-designing tokenizer, backbone, and system reaches competitive quality without scaling to tens of billions of parameters.

2026-07-27 · MONDAY · 00:40 PDT

  1. 1 owensong/Inflect-Nano-v2 252 DOWNLOADS · 79 LIKES The smaller sibling of Inflect-Micro-v2 does complete local text-to-waveform speech synthesis in under 4M parameters. It offers fixed-voice English output with deterministic seeds, long-text handling, and CPU or CUDA inference. The author says the release was built and funded independently.
  2. 2 Lightricks/LTX-2.3-22b-IC-LoRA-Clean-Plate 0 DOWNLOADS · 79 LIKES Lightricks published an in-context LoRA for its LTX-2.3 22B video model that produces clean plates, the VFX term for a shot with an object removed. It is tagged video-to-video and has drawn 79 likes with no downloads recorded yet.
  3. 3 poolside/Laguna-XS-2.1 23,776 DOWNLOADS · 181 LIKES Poolside released a 33B mixture-of-experts model with 3B parameters activated per token, aimed at agentic coding and long-horizon work on a local machine. The card reports a 5.4 percent gain on SWE-bench Multilingual over the previous Laguna XS.2, plus stronger terminal performance.
  4. 4 DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF 55,444 DOWNLOADS · 86 LIKES A community GGUF release of an uncensored Qwen3.5-9B fine-tune, shipped in both regular and MTP quantizations. The card claims the 4-bit and 8-bit builds beat all seven of its reference benchmarks for several larger Qwen models. It has drawn 55,444 downloads.
  5. 5 PaddlePaddle/HPD-Parsing 797 DOWNLOADS · 64 LIKES PaddlePaddle published HPD-Parsing, a hierarchical parallel document parsing model built on an InternVL chat backbone. The repo links an arXiv paper and vLLM serving, and is tagged image-text-to-text for visual-language document work.

2026-07-26 · SUNDAY · 12:37 PDT

  1. 1 owensong/Inflect-Micro-v2 47 DOWNLOADS · 116 LIKES A text-to-speech model that does complete local text-to-waveform synthesis in under 10M parameters. It offers fixed-voice English output with deterministic seeds, long-text handling, and either CPU or CUDA inference. The author says the release was built and funded independently.
  2. 2 microsoft/Fara1.5-27B 1,039 DOWNLOADS · 97 LIKES Microsoft released a 27B multimodal computer use agent for driving web interfaces from screenshots. The model card links a paper and an Azure Foundry deployment path, and the repo tags it as a CUA and web agent built on a Qwen3.5 backbone.
  3. 3 LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V5-GGUF 60,643 DOWNLOADS · 156 LIKES A community GGUF release of an uncensored, vision-capable fine-tune of the Qwen3.6-35B-A3B mixture-of-experts model. The card frames its Genesis method as removing accumulated tensor noise from training. It has drawn 60,643 downloads.
  4. 4 microsoft/Mage-Flow-Edit-Turbo 843 DOWNLOADS · 79 LIKES Microsoft published a 4B-scale rectified-flow stack for native-resolution text-to-image generation and instruction-based editing. The card argues that tokenizer, backbone, and system co-design reaches competitive quality without scaling to tens of billions of parameters.
  5. 5 unsloth/Ornith-1.0-35B-GGUF 90,764 DOWNLOADS · 99 LIKES GGUF quantizations of Ornith-1.0-35B, the MoE member of a self-improving open-source family for agentic coding. The family spans 9B and 31B dense plus 35B and 397B MoE variants post-trained on Gemma 4 and Qwen 3.5, with coding benchmark claims on Terminal-Bench 2.1 and SWE-bench.