site banner

Small-Scale Question Sunday for July 19, 2026

Do you have a dumb question that you're kind of embarrassed to ask in the main thread? Is there something you're just not sure about?

This is your opportunity to ask questions. No question too simple or too silly.

Culture war topics are accepted, and proposals for a better intro post are appreciated.

2
Jump in the discussion.

No email address required.

What are your views on Kimi? How seriously should we take benchmark claims, eg how rig-proof are the benchmark criteria which decide the performance comparisons of AI models?

I haven't been able to access the models on their free plan, which means I have no opinion. I've heard people I respect say it's a bit benchmark-maxxed, but the fuck would I know myself?

All I can say is that I don't see a particular use case for it. I have paid plans for ChatGPT and Max for Claude. Those have the bleeding edge models, and I always use the best I can get away with it. I rarely (but not never) need token-churning agentic work, where $/task is a strong constraint. I technically retain access to Gemini 3.1 Pro, but it is a sad model, and I'm not going to touch the family till they fix things with 3.5 Pro.

Chinese models are, in general, not very useful for me. I don't need them. I used R1 etc back when they did something useful for me, and I haven't been tempted back in the last half year.