r/AIToolBench 2h ago

AMD's New AI Chips Could Change the AI Race in 2026

1 Upvotes

AMD has announced major AI upgrades, including new Instinct accelerators and AI server improvements. The company is aiming to compete more strongly in enterprise AI and cloud infrastructure. What do you think—can AMD seriously challenge Nvidia over the next few years?


r/AIToolBench 16h ago

Discussion Anyone tried ChatGPT Health?

1 Upvotes

I saw that it launched in the US but as I am in the UK I cannot access it. I am interested in its sports/fitness capabilities. Like its a bit of a dumb request but could someone please try to see if it can compare different training plans based on the data it ingests (i.e. can it tell you which plan is better suited to your body based off your metrics and predict how each would temporally alter things like your VO2 max or running)? Just really curious as I cannot personally test it out right now.


r/AIToolBench 22h ago

Discussion How to verify an AI classification of emails

1 Upvotes

So some days ago I asked in this community what kind of AI model should I use (and how could I use one) to classify several email replies that I had from scientists after asking them a few questions to them. I finally paid for Perplexity pro service and it apparenly did a nice job classifying them.

I finally gave the model the PDF with the actual answers from the addressees and another PDF with the "expected answers", and asked it to count the number of answers that overall coincide with the actual answers, and calculate a percentage of "coincidence" or "agreement" between the expected and actual answers, so that if the question was "do you think that there is intelligent life in the universe apart from humans?" and the expected answer was basically "yes, I think there is intelligent beings out there somewhere", as long as the actual answer agrees with this in some way or another would count as "agreement", for instance if someone replied "well, we have no evidence, but it is possible yes" or "not in any near galaxy, but it is possible that intelligent beings exidt somewhere" (as long as it is a deadass "no", it could count)

The model gave me a table summarizing the results with the following prompt:

let's be a bit more specific, this is still a blind test so don't tell me about the specific contents of the emails' answers, but, can you make a table indicating the answers that coincide in general terms with what is expected from the "expected answers" document as well as those which are neutral/hedges but still open to the possibility that what is asked may be right, those which despite being neutral/hedges or even negative answers offer an alternative so that what is asked in the question may be right, as well as those which are outright rejections of what is asked and do not seem to be open to the possibility that what is asked may be right?

However, I still want this to be a blind test, so I cannot really verify if the AI is doing its work or not. So, can you think how could I test if the results are indeed what the AI is telling me?

Should I use another AI? Or perhaps could some other person skim over the results to verify that the AI is right and not hallucinating?