Shane Larson's picture

Shane Larson PRO

comgen42
1 11

AI & ML interests

None yet

Recent Activity

repliedto their post about 4 hours ago
Kodiak v0.2 1B got its official Decision Index listing this week: 11.8, rank 96 of 115. I'll be honest, that one stung. I've put a lot into this model. Our own run of the public benchmarks said about 19, and the index confirmed that part (19.1). But most of the full score comes from private tests, and on tasks from new domains Kodiak scored close to zero. It's good at what it was trained for: intent routing at 91.6, common sense, ranking. It falls apart on kinds of decisions it has never seen. Sarcasm on real tweets came out worse than random. There's a particular kind of tired that comes from working hard on something and having a scoreboard tell you you're near the bottom. Then I remember what I've learned from history. The Wright brothers came home from Kitty Hawk in 1901 convinced the published lift tables were wrong. Instead of quitting, they built their own wind tunnel and tested hundreds of wing shapes. James Dyson went through more than 5,000 prototypes before one worked. Failing wasn't the exception for the people who built things that mattered. It was the job. What separated them was that they kept going and kept measuring honestly. So that's the plan. The index just told us exactly where the weakness is: generalizing to new kinds of tasks. Meanwhile a Hugging Face user found two shortcuts in v0.3, we fixed one and are testing the fix for the other right now, and every number goes into the public log, good or bad. Rank 96 is a data point, not a verdict. Back to work. github.com/grizzlypeaksoftware/kodiak
posted an update about 15 hours ago
Kodiak v0.2 1B got its official Decision Index listing this week: 11.8, rank 96 of 115. I'll be honest, that one stung. I've put a lot into this model. Our own run of the public benchmarks said about 19, and the index confirmed that part (19.1). But most of the full score comes from private tests, and on tasks from new domains Kodiak scored close to zero. It's good at what it was trained for: intent routing at 91.6, common sense, ranking. It falls apart on kinds of decisions it has never seen. Sarcasm on real tweets came out worse than random. There's a particular kind of tired that comes from working hard on something and having a scoreboard tell you you're near the bottom. Then I remember what I've learned from history. The Wright brothers came home from Kitty Hawk in 1901 convinced the published lift tables were wrong. Instead of quitting, they built their own wind tunnel and tested hundreds of wing shapes. James Dyson went through more than 5,000 prototypes before one worked. Failing wasn't the exception for the people who built things that mattered. It was the job. What separated them was that they kept going and kept measuring honestly. So that's the plan. The index just told us exactly where the weakness is: generalizing to new kinds of tasks. Meanwhile a Hugging Face user found two shortcuts in v0.3, we fixed one and are testing the fix for the other right now, and every number goes into the public log, good or bad. Rank 96 is a data point, not a verdict. Back to work. github.com/grizzlypeaksoftware/kodiak
View all activity

Organizations

Cortex Agent LLC's profile picture