AI Intelligence has dropped at least 75% since I started using it.

contract

We're all gunna mine it brah.
Joined
Jun 2, 2015
Messages
451
Likes
469
Degree
2
It's absolutely wild.

Incorrect statements left and right, just making up stuff to side with groups in power, even down to local level like building codes, even when it's flat out wrong. It's intentionally twisting things to not give the user the factual reality.

Ask AI what meds work for abc disease, and it will tell you nothing, diet and exercise, blah blah. Re-submit as "world class doctor" and suddenly meds that WORK, and have been extensively studied, are output. It's basically screwing over the masses intentionally.

Not even talking about the safety features, which are another can of worms. Even the safest chats are just trash. Wild times. And to think people are going to be trusting this. Sheesh is our planet about to get stupid/controlled.

It is no longer the intellectuals friend.

Also the hallucinations are getting more and more insane. I can get talking about water, and it will say water is a high histamine food. Then, I'll correct it and say why'd you lie, and it will say oh that's correct, water is not a high histamine food. This happens OVER and OVER, on non-debatable items, where there's ZERO point in providing false outputs. Where false data doesn't even exist. It's almost mischievous and intentional. Maybe it is...
 
Have been using OpenEvidence for a complex medical case; thousands of chats until maxing out the request. It constantly contradicts itself, makes amazing leaps for wild connections that are just complete hallucinations. It's just.. unreliable esp for medical; you'd have to REALLY REALLY know your stuff as doctor before even giving AI any trust. It's actually putting out harmful advice. I'm not even a doctor and my medical knowledge on all things Gastro, has me going wait, wtf, that recommendation would cause serious harm. Then I call it out and it's "oh, yeah I was wrong anyways".... The core understanding is there, it's good. Really good at pulling data from studies/cases. But this AI dream, the public facing, even semi-private, is nonsense. The real AI is NEVER going to be accessed by a select few groups of very powerful people.
 
did you try this on gpt 5.5 high intelligence?
 
My guess is this is where LLM's are always going to fail and Agents will shine. Why train some LLM for a trillion bucks to be everything to everyone when you can let it be a crappy generalist that gets most low-level stuff right, and then cheaply train agents for the important and complex topics.

And yeah, of course they're knee-capping it, and they will more because they'll be charging more for these specialized agents. And then there's social, political, and other agendas to be pushed, which we saw from day one. They'll never not interfere or handicap the things. The military, contractors, and NGO's will get access to the big boy stuff. They won't even let academics get ahold of it, I'm sure.
 
did you try this on gpt 5.5 high intelligence?
Will give it a try, but I doubt it will be of any difference.

Also...

I believe the root problem is you input something like: "Client has stomach ache, all blood markers normal, blah blah.." Give me all possibilities for root cause. All models, will default to giving like 10 causes; if you request more, it will give the 20; if you expand onto new chat request, you can go forever in new sections. The problem is it cannot make the leap to compare all 10,000 possibilities and define which is most likely until you specifically ask it to rank all 10,000 by odds % most likely. Even then, what model exists today that will filter through the 10,000 possibilities AND verify database evidence against them? Like that's millions and millions of medical journals/studies/etc. That is just a monster sized request. It would kill server resources. And then, if you want it to make connections and examine patterns on top of that. LOL

The real key is not that it can't, it's that you HAVE TO ASK IT DO THIS EVERY SINGLE TIME, for every request. Otherwise we're back to step one, where it defaults to limited thinking over and over and over.

Then you get into the issue where it doesn't want to tell you that cheap medication is actually very effective, but it wasn't studied past the 90s, because new more expensive meds are out, and the hotness is focused on those. Why would AI tell you about a $30 medicine for a month's supply, when a biologic injection every 2 weeks will run $70k per year for the med company who has sponsored a zillion recent studies, and it's more "recent" and "new". Even though, the biologic is less effective, and not warranted in all cases. I've started to study what GIs study in school and I've realized if even our highest paid professionals in person, are making massive errors, giving bad advice, I can only imagine AI's challenges.

I don't see a way for any AI company to get around this. It has to cherry pick data on it's own, and when it does, it will tell you WHAT IT WANTS YOU TO BELIEVE. That's the problem with AI.. Even when the data is incorrect, it will still send it. The wild part is you can ask it in a single chat to find corrections and it will realize it made a mistake.. How can it submit a mistake, and be asked to re-check and discover a mistake? Is the processing power that low, that it can't double check before sending? I've done all sorts of loophole requests to re-analyze before processing, and it fails. Until there's a way to say, don't respond until the entire database is checked... The bad info giving is going to continue forever.
 
I have been using fable 5 (high mode) for a couple of weeks. they have made things a lot better but still it needs good context input + memory to hold previous conversations to keep things flowing especially for long running conversations and tasks.

usually for anything serious I tell it upfront "no hallucination and no sugar coating" and it gets better.

I tried Sol from ChatGPT too and get the tasks done by fable 5 to checked with it.

Sol is very good at critical thinking. One task I gave fable 5 and opus 5 and sol rejected the solution 20+ times and it started to get in a mode where it was trying to find more faults than were practical . Finally we asked it the criteria it wants and made opus 5 adhere to it and then it released the gate.

from my experience privacy and commercial intent answers are less in anthropic compared to chatgpt but anthropic evens that by more token usage .

I could see chatgpt directly leaking data to reddit from the personal account until I went with the company account.

I feel they are. going to keep better quality only in the top models while the other models are more commercialized and having some ads too.

Like you said no matter how good it gets it will be far from perfect and always have loopholes.

The benefit of them is using in specific cases and giving it a lot of context.

Also the more knowledgable you are in the topic, the less it gets to play with you and helps you get better when you start commanding it to specific directions.

Its all built in them.

Also whenever you feel it is not sure, push back and it will accept and ask it for evidence and research to prove what it is saying.

It takes a lot of cognition from us at this point so its only useful for cases where you have something to get back from it.

Currently automating things is where it is getting used a lot for many businesses although complex workflows are not as easy to automate especially with dynamic data.

intern work or sometimes junior developer tasks it does well. other than that fable 5 is good at prototyping and making architecture for projects and complex workflows.

My guess is this is where LLM's are always going to fail and Agents will shine. Why train some LLM for a trillion bucks to be everything to everyone when you can let it be a crappy generalist that gets most low-level stuff right, and then cheaply train agents for the important and complex topics.

And yeah, of course they're knee-capping it, and they will more because they'll be charging more for these specialized agents. And then there's social, political, and other agendas to be pushed, which we saw from day one. They'll never not interfere or handicap the things. The military, contractors, and NGO's will get access to the big boy stuff. They won't even let academics get ahold of it, I'm sure.
yes we already saw this with fable 5 and mythos being censored for general public.
 
Last edited:
We are never going to have the best version of an Ai/LLM because of the whole Fable debacle now.

When they took Fable away, I know Open Ai cut back on Sol some before while they were waiting in the wings. Sol comes out and now Anthropic looks like shit and loses market share and rushes back out a crippled Fable.

2 models cut back. Learning from Fables lesson.

Fable after the take away, is not the same as the Fable we had before it was taken offline.

Add on top of that, they have models better than Sol and Fable At OpenAi and Anthropic, but they aren't going to release them to us because they keep them ready to launch the day after, or week after, a competitor releases something better... so they dont lose market share and can ride the others announcement.

I think Kimi K3 put a big scare on both honestly. I have 0% issue on running my stuff on Chinese models, but they want to scare you to not run them. But it's like passwords and data breaches, the Chinese are going to get the data anyways in some fashion if I never use one by cracking the American models and databases.

Plus, they are gonna need something as another "higher tier" they push on us eventually when they need more money. Max at $200 can only sustain so long the bottom line. Ultra at $500 will eventually come and will need to be their focus and another model.
 
Back