Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Sol is so much better than Fable 5

I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Sol is a much smaller models and it shows. It often misses the forest for the trees.



I feel like a lot happened this week and people are glazing how ridiculously strong Flash 3.8 is right now compared to Fable/Opus/Sol/Astra.


Flash 3.8 is rad. Easily my daily driver now. Only downside is it's Gemini so sometimes it just keeps going until it wants to be done.


I have a few attention and finish mechanisms in my prompts. I have been using it for a week and a half and with some prompt taming it is great. (I have early access to the models cause I work at the place that makes the model). None of my attempts to ever tame Opus 5 have worked.


>> I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Same. It makes me wonder what types of things the person must be working on.


This is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it.

People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.


It’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.


This is the job now, we are shepherds.


This is why I generally don't trust benchmarks, or anything other than my own experience tbh. It always seems like everyone has a different answer.

If we truly had some AGI model, it would probably be fairly obvious to us all no?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: