Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

raincandy-uย 
posted an update 1 day ago
view post
Post
2387
20K parameters can tell a story. ๐Ÿš€

๐Ÿค— We trained a ~20k-parameter Transformer that can actually write stories!

raincandy-u/MacroStories

โ†’ ~50ร— smaller than the 1M-parameter TinyStories model
โ†’ ~3,000ร— smaller than AlexNet
โ†’ 81 KB in FP32

yayyy the whole model. เซฎ หถแต” แต• แต”หถ แƒ

She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block โ€” recurrently applied 4 times with shared weights.

Despite having only 19,969 parameters, she can maintain a simple narrative across 100โ€“300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.

She runs extremely fast on CPU โ€” no GPU required. The entire model is tiny enough to load almost instantly! โ˜บ๏ธ
  • 2 replies
ยท
Undi95ย 
posted an update 3 days ago
view post
Post
4131
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek.

I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use?

No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note.
The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match.
Later, code will be checked by actually running tests.

The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights.
Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again.

The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move.

I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks.

Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try.

Did you already tried something like that? What was your result? I'm curious!
  • 21 replies
ยท
SeaWolf-AIย 
posted an update about 11 hours ago
view post
Post
977
๐Ÿง  We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.

Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.

โš™๏ธ How it works
๐Ÿ”น It makes its call in a single forward pass.
๐Ÿ”น Zero generated tokens, and no decoding loop.
๐Ÿ”น That keeps latency and cost far below what a generative model needs.

๐ŸŽฏ What it judges
๐Ÿ”น It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
๐Ÿ”น For each one it hands back a calibrated confidence, not just an answer.

๐Ÿ“Š How well calibrated (measured)
๐Ÿ”น KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
๐Ÿ”น 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
๐Ÿ”น By type: noul 0.847, choice 0.723, score 0.675.
๐Ÿ”น None of the benchmark's train split went into it. It is pure zero-shot.

๐Ÿš€ Where it fits
๐Ÿ”น Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.

๐Ÿ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).

๐Ÿ”— Links
Model: FINAL-Bench/Darwin-27B-ZTC
Leaderboard: LocalLLaMA/typed-decisions

Curious to hear what you make of the single-pass, no-generation approach. ๐Ÿ™Œ
  • 2 replies
ยท
danielhanchenย 
posted an update 2 days ago
view post
Post
3498
Google releases EmbeddingGemma 2, a new open embedding model that runs locally on 0.5GB RAM.

The 740M parameter Apache 2.0 model combines a 270M text model with vision (170M) + audio (300M).

Run & train the model via Unsloth.

GGUF: unsloth/embeddinggemma-2-GGUF
Guide: https://unsloth.ai/docs/models/embeddinggemma-2
Parveshiiiiย 
posted an update 1 day ago
view post
Post
2769
Most deepfake audio detectors are quietly cheating.

They donโ€™t really listen to the speech โ€” they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.

AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.

The result is a detector that actually has to learn the artifacts instead of gaming the feature space.

Model: Modotte/AIRealNet-Audio
DedeProGamesย 
posted an update 2 days ago
view post
Post
3870
Im working on a 23M ASR model, trained on 100k hours of audio
  • 4 replies
ยท
comgen42ย 
posted an update 2 days ago
view post
Post
3137
๐Ÿป Kodiak-v0.3-1B: seven new kinds of decision, measured on real data.

Kodiak is an open 1B encoder that answers typed questions with calibrated confidence, or says "can't tell". New in v0.3:
โ€ข pairwise judge (which answer is better?)
โ€ข long-answer hallucination checks
โ€ข stance and sarcasm
โ€ข policy violation and refund-eligibility checks against your written rules
โ€ข agent step safety (run it, ask first, or never)

We stopped trusting our own synthetic tests and judged v0.3 on real labelled data it never trained on (RAGBench, MT-Bench human judgments, SemEval stance): 0.21 โ†’ 0.36, averaged over 3 training runs. Nothing else got worse, and when it says "can't tell" it's right 93% of the time (up from 88%).

Known limits are in the model card.

๐Ÿ“ฆ cortex-agent-llc/kodiak-v0.3-1b
๐ŸŽฏ Accuracy mode: cortex-agent-llc/kodiak-v0.3-1b-accuracy
๐Ÿ•น๏ธ Demo: comgen42/kodiak-demo
๐Ÿ“ What we built, and the kill that changed our process: https://cortexagent.com/blog/kodiak-v0-3-seven-new-kinds-of-decision-measured-on-real-data
  • 4 replies
ยท
CountingSheepย 
posted an update 3 days ago
view post
Post
1715
Vev: Jev-style decisions about images, from a 4B or 9B model you can run yourself.

Give it a screenshot or photo plus a yes/no, multiple-choice, or scoring question, and it returns a probability for every option instead of generating text.

It serves TypeSafe's /v1/systemone format, so the official SDK works by changing the base URL, including requests with images.

The clip shows vev-4b playing Doom in real time. On each look, the image is split into 8 vertical slices, and Vev answers 8 yes/no questions in a single request, one per slice. A small fixed-rule harness turns those probabilities into turning and firing.

Try it in the browser:
CountingSheep/vev

Weights:
CountingSheep/vev-4b
CountingSheep/vev-9b

LoRA adapters are also available in the collection.

Code:
https://github.com/Xiaooolong/vev

Fine-tuned from Qwen3.5. Tested on NVIDIA GPUs so far.
tardellirsย 
posted an update about 24 hours ago
view post
Post
1479
Robotics is now the second most downloaded dataset category on the Hub, after text generation.

Robotics datasets got 13.7M downloads in September, ahead of text classification and question answering. Two years ago the category ranked 23rd. One in 7 new datasets is now robotics, mostly LeRobot recordings: typically a few dozen demos, about half of them on low-cost SO-100/SO-101 arms.

I found this after adding datasets and Spaces to Model Pulse, which rebuilds daily history from @cfahlgren1 's hub-stats snapshots. Two more findings:

- In 2022, 36% of authors who list training data cited classic NLP sets like IMDb, SQuAD and GLUE. In 2026 it's 2.4%. Reasoning traces distilled from models like DeepSeek-R1 and Claude are now the most cited kind.
- In October 2025, 122K Spaces were created, 71K of them websites built with DeepSite. That's about 6x the monthly pace of late 2024, while likes given per month fell from about 35K to about 20K.

New in the app: a page for every dataset, with daily downloads and the models trained on it (636 list FineWeb), a page for every Space, and rankings for both.

Thank you to everyone who liked Model Pulse this week: it made Spaces of the Week and is #7 on trending. Thanks also to @dipankarsarkar , whose comments on the last post fixed three data issues. If a number looks wrong, tell me.

Spaces: tardellirs/model-pulse
  • 5 replies
ยท
prithivMLmodsย 
posted an update 1 day ago
view post
Post
1501
OneDecision-VisionGuard-Demo is now available on Hugging Face Spaces!

๐Ÿค— Space: prithivMLmods/OneDecision-VisionGuard-Demo

This demo showcases the OneDecision-VisionGuard family of multimodal image classification models for detecting NSFW and other sensitive visual content, with structured JSON reasoning, improved accuracy, and better handling of edge cases such as sensitive imagery, uncensored analysis, scene descriptions, and classification reasoning.

๐Ÿ“ฆ Models: 27B, 9B, 4B โ€” prithivMLmods/OneDecision-VisionGuard-27B-SFT, prithivMLmods/OneDecision-VisionGuard-9B-SFT, prithivMLmods/OneDecision-VisionGuard-4B-SFT

โ†—๏ธ Collection: https://huggingface.co/collections/prithivMLmods/onedecision-visionguard

To learn more, visit the app page or the respective model pages.