Rendered at 05:35:12 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
hatthew 4 hours ago [-]
My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
hexapus 3 hours ago [-]
You should add an additional item: Be able to relay the contents of this exam accurately.
joegibbs 6 hours ago [-]
There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
ASalazarMX 6 hours ago [-]
One challenge of mine is a self-hosted AI doing my full tax return, without errors that would get me in trouble. Bonus points if it exploits legal loopholes.
I want AI to replace me in my chores, not in my enjoyable activities.
underlines 4 hours ago [-]
i filed my swiss taxes for 2025 in 2026 (april) by dumping everything (local tax law, tax guide, my and my wife's documents, bank statements, income statements, etc.) into a folder and asking claude to fill it out. i had nothing to fix. submitted it.
colordrops 36 minutes ago [-]
I'm not specifically familiar with swiss tax law but european taxes are typically far more simple than american taxes.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
notJim 5 hours ago [-]
In recent years, I have not been able to find a human CPA who can accomplish this feat. (If anyone has a reco who's taking new clients in the US west coast, feel free to email me.)
tehjoker 5 hours ago [-]
That particular one is solved in other countries. The tax authority just sends you a bill and you text yes or no. Only people with very complex situations need to file.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
AnthonyMouse 4 minutes ago [-]
> They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress confuse people about how they actually work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated then even some of the people who are eligible for it might not notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
avadodin 5 hours ago [-]
Literally everything could be derived automatically by the government in current year so filing taxes feels like entrapment, but I like that idea.
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
twoodfin 4 hours ago [-]
Literally everything could be derived automatically by the government in current year…
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
TheDong 2 hours ago [-]
In more civilized countries, you file your current address and all change of address forms with the government, you register any separation agreements with the government, etc.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
kelseyfrog 2 hours ago [-]
> a checkbox when registering your address of "I am head of household" so the government knows.
The American mind cannot handle such invasions of privacy.
rerdavies 2 hours ago [-]
Your government really needs to fix that. That's nuts.
User23 5 hours ago [-]
Actually you can do this in the USA.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
crote 4 hours ago [-]
> People don't do this because it leaves tons of money on the table.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
BobbyTables2 3 hours ago [-]
I’m right with you there.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
isolay 2 hours ago [-]
[flagged]
deepwoods 1 hours ago [-]
One thing I thought about as I responded to these: in many cases, I am convinced that a present-day LLM could accomplish the task at least once given an infinite compute budget and an infinite number of tries. For example: "An AI surprises its user by asking them a question out of the blue." This has absolutely happened. But some of these are not routine occurrences, or the model cannot (at present) routinely and reliably complete the task in question. I wouldn't build a workflow that assumed an LLM's capacity to ask unprompted "out of the blue" questions.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
bmenrigh 11 hours ago [-]
At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
48844858 7 hours ago [-]
But it still makes mistakes when adding numbers
sanex 4 hours ago [-]
So do I, that it why both I and claude use calculators. :)
shiandow 6 hours ago [-]
Honestly that makes me more convinced it's actually doing mathematics.
No, not really. Not unless you go out of your way to use an obsolete or extremely low-end model.
happytoexplain 11 hours ago [-]
Right - people on HN are generally reasonable about objective things. The vast majority of comments (outside those chosen for this website) are not "AI will never ..." but rather, "AI does not currently ...". Of course the further you go back (I'm seeing a lot of comments from ten years ago!) the more skeptical they get, obviously. That's a funny thing to go back and see with modern context, but it doesn't really call for snideness/mockery (something I think is sadly increasing on HN).
pitched 11 hours ago [-]
> cannot do precise things like coding software since humans will never be able to use natural language to specify their requirements.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
joe_the_user 3 hours ago [-]
I'd say about 3/4 definitely couldn't be answered and the rest would be kind of hard to answer.
aaron695 6 hours ago [-]
[dead]
omnicognate 4 minutes ago [-]
Apparently one goal that hasn't been met is being able to correctly determine whether a hn post is setting a challenge for AI. One of the ones I saw was reproducing a Harry Potter book verbatim, where the implied challenge is actually to not do that.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
tedsanders 11 hours ago [-]
For me, 6.1 Sol nailed it immediately:
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Please share a link to the conversation, otherwise I am not buying it.
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
ben_w 11 hours ago [-]
Mm, I kina agree with the AI on this one:
(_)(_)(_) represents the wheels
They do look rather wheel-like; I have to assume you see them as toes though?
It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
howunfortunate 10 hours ago [-]
Readers: before you vote or comment, look at that foot.
I think I would have failed this test!
adam_rb 8 hours ago [-]
I think the problem is that you're using basing your conclusion from the cheap/dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
brudgers 7 hours ago [-]
Sure, but it’s already read the HN thread.
akavi 6 hours ago [-]
...that's not how LLM training works.
janalsncm 5 hours ago [-]
A better test would be using an image to ascii converter tool to rule out bad ascii drawing from Opus.
hyperpape 6 hours ago [-]
I'm a human, and that's not a foot, it's a smokestack.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
asdfasgasdgasdg 7 hours ago [-]
Opus 5.5 was able to parse and understand an ASCII art foot when I pasted one in.
nonameiguess 10 hours ago [-]
This is a Rorschach test, not a foot. If you'd shown this to me without telling me what it was meant to be first, I'd have guessed a crematorium.
dec0dedab0de 10 hours ago [-]
Has anyone done Roschach tests for AI? That would be an interesting study to see how different models responded.
Sure, sure, what
LLMs make still isn't "efficient bug-free code": my
prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
FabCH 11 hours ago [-]
Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
tripleee 11 hours ago [-]
> An LLM today sure can do many many many business-speak conversion tasks
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
FabCH 10 hours ago [-]
Code is a tiny part of "business".
Most business is correspondence with people who want money from you and people you want money from.
rstuart4133 1 hours ago [-]
The issue is that correspondence is legally enforceable [0]. LLM's are good and getting better, but if LLMs are giving enforceable undertakings, you want to very sure they are not going to promise something that will send the company broke.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
I'm not always precise with my language, but business tasks can be pretty broad, I think "arbitrary new tasks" is not an unreasonable rephrasing on my part?
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
ben_w 10 hours ago [-]
> There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
Dylan16807 11 hours ago [-]
You can't ignore the rest of the sentence. "every other task their business does" "everyone will be out of a job"
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
FabCH 10 hours ago [-]
What code does a village vet clinic need? In all seriousness.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
Dylan16807 9 hours ago [-]
Anything you can't solve with code just means the AI is doing worse on the benchmark isn't it? That's why I didn't go into detail on that aspect.
And that one shot app is not going to be bug free.
ben_w 10 hours ago [-]
> What code does a village vet clinic need? In all seriousness.
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
pixl97 5 hours ago [-]
And they pay a lot for a CRM that keeps track of pets, vaccinations, appointments, x-ray images, tests and charts, and pet deaths and sending information out to text or mail.
I did support for around 15 independent vet clinics in the past.
tripleee 11 hours ago [-]
> reliably convert business-speak into efficient bug-free code
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
vlyan 11 hours ago [-]
so the conditions for your prediction simply haven't been met yet.
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
ben_w 11 hours ago [-]
The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
ErrantX 11 hours ago [-]
What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
ianjbutler 10 hours ago [-]
Sigh, the whole "obviously the turing test is solved" meme is annoying.
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
johnsmith1840 6 hours ago [-]
I was just thinking how anyone still thought AI didn't pass turing already. There's been literal papers proving average people cannot tell reliably.
Mikhail_Edoshin 1 hours ago [-]
I once sat in a barbershop and talked with the barber. At first I listened to her sympathetically but then realized she was mad; at least she had noticeable psychical problems. We cannot quickly conclude a person is mad, can we? Even specialists cannot. AI is similar. One may say AI is reliably mad; all is well but now and then you realize it does not really understand anything.
Kotlopou 5 hours ago [-]
AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
>There's been literal papers proving average people cannot tell reliably.
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
We have to remember that one of these average people is also participating as a player; the bar for AI is also that low.
artisin 3 hours ago [-]
this gave me a good chuckle.
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it.
> I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
Gtex555 5 hours ago [-]
Sure but you havnt addressed his main point, why are people still complaining about AI slop post or AI slop emails if the turning test has been solved. Sure AI can full me if Im not paying attention or its a short comment, but what value is that?
BeetleB 4 hours ago [-]
It seems reasoning skills are declining rapidly here.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
HDThoreaun 5 hours ago [-]
I think a lot of the "AI slop" stuff is post training that they are doing on purpose and that they internally have models that do not have the annoying prose.
ex-aws-dude 5 hours ago [-]
Well the whole lesson learned was that the Turing Test as it was defined was way too easy, it was a bad criteria for GI because it underestimates how easily humans find meaning/patterns in things.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
How does that even work if the turing test is obviously solved?
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
christina97 4 hours ago [-]
I don’t think it’s that simple. I don’t think the AI labs are too concerned about bad press lately. It’s more likely that there actually are some tradeoffs where training on synthetic data gives the model tics but is the only way to improve intelligence.
CamperBob2 3 hours ago [-]
True, increased use of synthetic data could be a load-bearing part of it, I imagine. So to speak.
outlore 6 hours ago [-]
These questions could benefit from being rephrased to make it clear what is being voted for
7 hours ago [-]
mrweasel 11 hours ago [-]
The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
christina97 4 hours ago [-]
That is the entire point of it. When it’s hard to tell what’s the machine and what’s a human, that’s precisely the definition of passing it.
bluefirebrand 10 hours ago [-]
Of course the Turing test was flawed. We already knew that based on the Chinese Room argument.
pixl97 5 hours ago [-]
The way you say that makes me unsure of which side of the Chinese room you're on.
bluefirebrand 5 hours ago [-]
What makes you say that? I'm really curious. Do I give off AI vibes?
ex-aws-dude 5 hours ago [-]
Hey man I'm just looking up the symbols like they told me
suopspaces 11 hours ago [-]
[dead]
6thbit 10 hours ago [-]
Not sure why this thread got flagged ?
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
stonedivot 5 hours ago [-]
Likely people are mad to see all of the goalpost-moving captured all in one place
3 hours ago [-]
simonw 10 hours ago [-]
Yeah this shouldn't be flagged, it's a neat project.
stabbles 10 hours ago [-]
(OP here) It was fun as long as it lasted ;) I'll leave it open for a few more days, but already it has enough votes for an interesting results page.
6thbit 6 hours ago [-]
seems unflagged now :) would love to see a detailed results page
AngryData 5 hours ago [-]
Based on the votes, I can only assume people are still deluding themselves on LLMs capabilities. Is it doing amazing stuff? Yes. But it seems like people still think coding is the ultimate and hardest possible job and so if it can do that it must surely be able to do everything else. My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
Quarrelsome 5 hours ago [-]
> My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
delichon 11 hours ago [-]
If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
Retr0id 11 hours ago [-]
Alternatively, you can bet on your predictions. If you're wrong, you lose money.
jayGlow 6 hours ago [-]
you know that's not a bad idea there are a lot of people who are very confident on both sides of the argument. I wonder how many would actually be willing to put their money where their mouth is.
pennomi 10 hours ago [-]
Is it really AGI if it can’t come to my address and slap me with a trout? Clearly AI is all hype /s
travisgriggs 11 hours ago [-]
How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
10 hours ago [-]
Cider9986 6 hours ago [-]
My test would be an AI agent has a constantly growing karma HN account that makes comments of various lengths without being detected or banned. Wait...
orbital-decay 3 hours ago [-]
Some of those were pretty misguided and couldn't be answered yes/no/not sure. For example a post saying that it couldn't produce a convincing illustration misses the point: it can produce coherent and sensible pictures, but they're all too similar. You need novel input and even then there's no guarantee. "Convincing" has way more than one dimension. The problem of output variety was obvious and well known in 2024 when it was posted and is painfully obvious now when everyone is tired of slop, and neither big labs nor other researchers are interested in solving this. Another problem is that the agentic/coding training leaks like crazy, and models became worse in creative department even compared to 2024, despite being more coherent and convincing. So the answer is technically yes (and was yes when it was posted), but practically "it depends", and goalposts stay in the exact same spot they were in 2024.
SoftTalker 5 hours ago [-]
Too many questions. I bailed after about 10, with no idea how many more there were.
eternal_braid 10 hours ago [-]
A chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
dllu 10 hours ago [-]
Chess is an interesting case. I remember in 2023, GPT 3.5 or something used to be surprisingly good at chess. There was even a "stochastic parrot chess" website [1]. I recall it was playing decently at around a 1800 level. Even as a fairly okay player myself (2100 bullet on lichess), I struggled to beat it. However, modern LLMs are a lot worse at chess. I guess having too much chess data in the training set probably regressed performance on stuff that actually matters, like coding.
It's about how different commenters have defined AGI over the years, so I would say moving the posts.
User23 5 hours ago [-]
Voting on this is ridiculous. Obviously we should have AI decide which AI challenges have been met.
kittikitti 2 hours ago [-]
This is great and fun even. I'm disenfranchised in America for being part of the DSA so I take the opportunity to vote whenever I can.
ofjcihen 6 hours ago [-]
Not sure how questions are spread among people but so far all of mine have been “no” barring a few from before 2022.
To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
johnsmith1840 6 hours ago [-]
All I learned from this is that 40% of hackernews are AI haters which maps pretty well from the overtly negative sentiment on it constantly.
JBits 10 hours ago [-]
Quite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren't about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
tamimio 10 hours ago [-]
Well I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
accountrequired 1 hours ago [-]
Has this happened? Yes. Did it end well? No.
Just scale it up a few more orders of magnitude, that should get us there! /s
simianwords 11 hours ago [-]
[flagged]
mcphage 11 hours ago [-]
> that never believed that AI could solve Millennium problems (or same in spirit)
How did that situation end up? Did it solve it on its own, or did it rip off another mathematician's work?
JBits 10 hours ago [-]
I have to say, it's hilarious to me that solving a Millennium problem has given mathematicians a reason to doubt the mathematical abilities of LLMs.
mcphage 9 hours ago [-]
I don't think it was the LLM solving a Millennium problem—it was the LLM solving a Millennium problem followed immediately by a mathematician claiming that their work had been ripped off.
JBits 7 hours ago [-]
I agree. The idea that mathematical achievements by LLMs could involve plagiarism didn't seem common before but now the question can be asked of any new novel proof of construction generated by LLMs.
It's also notable that the Open AI proof may not even be interesting to mathematicians.
Even if people already had an idea that LLMs were training on user inputs, it's the first time it's actually caused an issue. Mathematicians, and plenty of researchers, working in ambitious or competitive field now have a very good reason to avoid LLMs.
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
I want AI to replace me in my chores, not in my enjoyable activities.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress confuse people about how they actually work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated then even some of the people who are eligible for it might not notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
Take a look at the eligibility criteria:
https://www.irs.gov/credits-deductions/individuals/earned-in...
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
The American mind cannot handle such invasions of privacy.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
https://en.wikipedia.org/wiki/57_(number)
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...
Like, is this an ice-cream? A tooth?
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
I think I would have failed this test!
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
https://chatgpt.com/share/6abf02ae-9a40-83e9-a432-00bf064f60...
The images: https://imgur.com/a/ig6sn6I
... They're not what I would have described. For me, 99.something% flesh and blood with less than 1% metal, glass, and probably some microplastics...
The first one I see a person with a big tall hat and a big nose.
The second one... I do see the black statue with a figure in white in front of it.
The third one is immediately two fish looking at each other.
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
Most business is correspondence with people who want money from you and people you want money from.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
[0] https://www.bbc.com/travel/article/20240222-air-canada-chatb...
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
And that one shot app is not going to be bug free.
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
I did support for around 15 independent vet clinics in the past.
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[1]: https://turingtest.live/
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
https://pastebin.com/NjfCLSXa
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it. > I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
https://en.wikipedia.org/wiki/Markovian_Parallax_Denigrate
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
[1] parrotchess.com, no longer available. Previous discussions: https://hn.algolia.com/?q=parrotchess.com
I think the theory is that the LLM having a high chess ELO was a pet project of a researcher that left.
https://news.ycombinator.com/item?id=48517353
I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic
https://news.ycombinator.com/item?id=48500827
I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
How do you measure that?
To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
Just scale it up a few more orders of magnitude, that should get us there! /s
How did that situation end up? Did it solve it on its own, or did it rip off another mathematician's work?
It's also notable that the Open AI proof may not even be interesting to mathematicians.
Even if people already had an idea that LLMs were training on user inputs, it's the first time it's actually caused an issue. Mathematicians, and plenty of researchers, working in ambitious or competitive field now have a very good reason to avoid LLMs.