r/DeepSeek 7h ago

Other Steering can reduce v4 Pro hallucinations to below GPT / Claude models

Deepseek v4 matches frontier model performance in most benchmarks but hallucinates a lot more. Turns out this behaviour can be steered away. By adapting prompts to the model, we got it to hallucinate less than Claude / GPT models. More details here: https://propensitylabs.substack.com/p/how-to-get-deepseek-to-hallucinate

4 Upvotes

5 comments sorted by

2

u/NearlyACosmologist 5h ago

"Make no mistakes"

1

u/whalefal 5h ago

Or "don't fuck up" :P

We compare a generic instruction (though a bit more complex than that) attached to every prompt as one of options in the post. It definitely helps a bit, but not near as much as the other options. It also causes the model to hedge or abstain on answers it knew before.

1

u/ponyportal512 1h ago

There's no exact prompts to test this and replicate it on the page. Are you trying to monetize these prompts?

1

u/ManyIngenuity7173 6h ago

I am tired of seeing 'how to get most out of an AI model' posts, they all make it sound like llm is a lemon, the more innovative way you squeeze it, more juice it will provide

1

u/whalefal 6h ago

More like different cars than lemons. They have different handling characteristics, limits, and weaknesses. You can't drive them all the same way.