I would really love if we brought back some colloquialisms in this field. Not that long ago most folks in tech would have had pretty blank looks on their faces when someone started talking about the "Pareto frontier"
> Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
My Kia Niro EV, which I'm otherwise quite happy with, has an occasional failure mode where it locks out the ignition for an arbitrary amount of time (generally 15 minutes). It's sometimes, but not exclusively, triggered by scheduled departure turning on the AC. Kia's response is a big shrug...
Think they're referring to the following, when Amodei was still working for OpenAI:
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
I was curious what he was responding to. Per https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff... it was "On July 19, 2019, McCandlish wrote in an OpenAI Slack channel: “We’re not sure if we’re going to release the Foresight LM Scaling paper publicly, but if we do we were thinking about removing all mentions of LibGen, since it's a bit of a sketchy data source." The paper may or may not be https://arxiv.org/abs/2001.08361 where they write "we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-available Internet Books."
Dario, if you’re reading this, I want you to know that Claude sucks now. Its output isn’t even English anymore, it’s just claudeslop.
You’ve actually ruined it so much that it has now polluted the Chinese models which are copying your work. So now all the models output claudeslop.
You’ve polluted all of the training data in the world. Now we’re never going to be able to train proper models, because everything has slop in it and all the sites have locked down their data to prevent future startups training.
I’m not the person you should say that to, I already don’t use them. My previous comment was specifically about “report that to their customer support”. Anthropic doesn’t have actual customer support
I would to order two soft taco supremes, a beef chalupa, and some cinnamon crispas. I'm sorry, what? Cinnamon crispas are discontinued? Since when? 1988! Unbelievable!
It's kind of a pain that you still have to file US taxes in this case. You may well end up paying zero in the US, but it's up to you to provide the necessary paperwork to show that your overseas income tax was high enough to cancel out the US income tax.
Even as someone who lived a long time in the US, I've failed a bunch of these on unclear definitions.
When you say "stoplights", is that just the lights themselves, or also the poles that support them? Is a tile that contains 1/3 of a bike wheel "bicycle" or "not bicycle"? Is a coach also considered a bus?
The problem with recaptcha style systems is that once the underlying model is trained, you are basically playing family feud - you aren't trying to provide a correct answer, as much as you are trying to match the "survey says..." majority opinion
I actually found I had more success with reCAPTCHA when I made myself think less about what the answer was.
I also learned this trick that a lot of the questions have a standard number of answers – e.g. there are always three buses, or fire hydrants, or whatever – so once I've picked three, I'm done. (There often is a fourth ambiguous case, which it accepts but doesn't require.)
I think I'm failing the captcha by not being sufficiently pedantic. It genuinely never occurred to me that a coach might not be considered a bus by the algorithm (where "algorithm" is presumably "distilled responses of many thousands of other humans")
Outside of the US, are there any markets BYD can't sell that car into? The Dolphin is currently listing at well under €20k here in Spain, which means its only real competition is the equivalently incredibly cheap Dacia Spring
I'm sure, but the Dacia brand as whole has been here for a long time, so they should have a reputational advantage over a newcomer like BYD. The Sandero compact SUV in particular is extremely popular (in a large part due to how cheap it is)
I think it's worse than that. Or at least, the products that have been built because of advertising are worse than that. Instagram/TikTok/Facebook are all ultimately just surfaces to serve ads, but they've wormed their way into their users' consciousness in fairly insidious ways
reply