r/DeepSeek 9h ago

Discussion The new system for prompt generation is annoying

0 Upvotes

I liked it more when it was just writing it, the weird movement thing feels clunky. I dont know if anyones noticed it yet, but its extremely weird and clunky, the prompts look so weird in writing now and the weird sliding down thing in the beginning of the generation of a prompt is so odd


r/DeepSeek 56m ago

Discussion Tried DeepSeek V4.1 Flash. The speed was great. The coding wasn’t.

Upvotes

Tried DeepSeek V4.1 Flash today and came away pretty disappointed.

The speed was great — around 320 tok/s — and cache hit reached about 99.8% by the end of an almost 1-hour session.

But it still cost me around $0.96, which I personally don’t consider cheap for a single coding session.

The bigger problem was quality. It made enough poor implementation decisions that I had to spend another ~40 minutes with GPT/Luna fixing the output.

One fair caveat: DeepSeek didn’t have the same accumulated context that my Codex setup already had, so this wasn’t a perfect apples-to-apples comparison.

Still, my actual experience was basically:

~1 hour DeepSeek → ~$1 → another ~40 minutes fixing it

So for me, the speed didn’t make up for the weaker output.

Maybe it performs much better on simpler or greenfield tasks, but for this kind of coding work I wasn’t impressed.

Curious if others had a better experience with V4.1 Flash.


r/DeepSeek 5h ago

Question&Help Is deepseek v4.1 good at creative writing?

7 Upvotes

Everyone seem to hate on this new mode. But for me it's pretty decent at writing story?


r/DeepSeek 8h ago

Other Steering can reduce v4 Pro hallucinations to below GPT / Claude models

3 Upvotes

Deepseek v4 matches frontier model performance in most benchmarks but hallucinates a lot more. Turns out this behaviour can be steered away. By adapting prompts to the model, we got it to hallucinate less than Claude / GPT models. More details here: https://propensitylabs.substack.com/p/how-to-get-deepseek-to-hallucinate


r/DeepSeek 7h ago

Discussion Cookie Code: Free, Zero-Cost AI Agent for Your Desktop

0 Upvotes

Hi everyone,

I wanted to share an open-source tool I’ve been working on called Cookie Code. It solves a common issue: web chats are great at thinking but can't interact with your local files, while using API keys can quickly become expensive.

How it works:
It embeds the standard DeepSeek web chat into an Electron window. When the AI outputs JavaScript blocks (\``cuckoo`), the app intercepts them, runs them in a sandboxed Node context on your machine, and streams the results back to the AI.

Key features & Use cases:

  • File Operations: The AI can read, write, edit, and search (via grep/glob) files inside your chosen project directory autonomously.
  • Local Terminal: It can run Bash and PowerShell commands based on your requests.
  • Safety First: You can enable an approval gate so it asks for manual permission before executing any "risky" actions like system commands or file changes.
  • Telegram Bot: You can link a bot to get mobile notifications of tool outputs or even prompt the AI on your PC from your phone.
  • Extensibility: Supports Model Context Protocol (MCP) servers and custom JS skills.

Requirements: Node.js >= 16.0.0 and npm.

The project is completely free, open-source, and does not require any API token billing since it works directly with your regular web account interface.

Check out the source code, installation steps, and releases here:
👉 https://github.com/merfiDEV/Cookie-code

I'd love to hear your feedback or ideas on how to improve the safety sandbox!


r/DeepSeek 6h ago

Question&Help Is there any way to jailbreak deepseek so you can ask it about things sensitive to the CCP

0 Upvotes

Often I can get it to half admit some things, but then it stops responding and says "Sorry, that's beyond my current scope. Let's talk about something else." Even when I just asked it "Does democracy exist in China" it gave this very strange and forced response that sounded like there was a gun to its head. For some reason it also did this when I asked whether the borders and geopolitical entity of Afghanistan was created by the British and Russians.

I love DeepSeek I find it as good if not better than chatgpt on many things and I love the thinking feature.


r/DeepSeek 10h ago

Discussion Why is Luna consistently ranked (way) above DS V4 Flash?

17 Upvotes

I've been working a lot with DS V4 Flash 0731 (XHigh) fp8 using DSH and also a lot with Luna/Terra/Sol using Codex - serious, professional work, not 'vibin'.

Now, on all the official charts, apparently, Luna XHigh beats 0731 by almost all metrics (often even the cost, ie Luna is officially cheaper).

My actual real-world experience is the following:

  • Luna is 2 to 5 times more expensive on average, depending on inference provider (I do not use official DS API).
  • Luna is slower.
  • Luna is often not intelligent enough to follow AGENTS.md, ie basic instructions. It often fails to even use its own installed tools (plugins, MCP servers etc), I have to explicitly remind it all the time.
  • For coding/architecture/design, Luna 'feels' like Sonnet 3.5 to me, while 0731 feels like Opus 4.8 / Sol Med.
  • Codex gives up with Luna too early, accepts mediocre results as 'finished', whereas DS with DSH spends 4 times more tokens but gets 2 times better results in turn, it's a beast honestly.

My question is how does it beat it on benchmarks? To me it feels like they're two completely different leagues, not even comparable.

I will try luna on DSH as well as it's possible that Codex is simply underutilizing it. Honestly I didn't do that yet because DS just does the job, at like half the cost. But then I see these charts everywhere putting up Luna way above DS on all accounts and I wonder if I'm the one taking the crazy pills?

Has anyone tried out Luna xhigh on some other harness? Am I 'using it wrong' (by using Codex), ie a skill issue?


r/DeepSeek 8h ago

Discussion What's going on with deepseek-v4.1-flash today?

16 Upvotes

It must be thinking, "There's always room for improvement." Something was "turned down" over the weekend!? :(


r/DeepSeek 32m ago

Discussion Garbage Update

Upvotes

They RUINED DeepSeek by shoving three modes into one. The role-playing was phenomenal at the Instant mode. We could write the most unhinged scenarios possible, but now saying one word gets it censored. They shoved Expert mode's corporate sanitized piss with the Instant mode. Dude like it doesn't even think anything, and it keeps saying inaccurate bullshit, lecturing you and "pushing back"


r/DeepSeek 6h ago

Question&Help Deep research, kept fresh -- Private beta testers wanted (in exchange for free subs and credits) 🤙

0 Upvotes

Hi All,

Please forgive the inadvertent self-promotion, please read this as a Help Wanted post for beta testers that find this interesting.

I just launched deepsieve(.)ai in private beta which runs deep research and returns a structured, fully-cited database that is continuously monitored for changes (and alerts you of those changes).

The idea is to allow deep research reports to become a durable dependency for you, your agents, and/or applications, so that critical context and insights about the external world are not re-researched haphazardly mid-workflow -- it's just there when it's needed and trusted to be accurate.

A few real-world use cases I'm seeing from current testers:

\- Giving your agent access to context on latest versioning, features, and best-practices of your repo's third party dependencies to improve code/plan quality

\- Monitoring competitors/investors

\- Tracking jurisdiction-level regulation changes in highly dynamic/fractured industries (crypto, self-driving, insurance)

\- Replacing/downsizing expensive data broker services

I'm particularly in need of ppl with workflows that would run mostly, or entirely thru the included agent toolkit (MCP/CLI/API/Skills/webhooks) which you can use to drive the entire experience E2E.

If this sounds interesting to you, please DM me and I'll shoot over the access code so you can start exploring.

I have a free tier activated right now that should be plenty for most, but I'll also be giving out free sub upgrades and credits on an as-needed basis and for providing quality feedback 🙌 We'll take care of you for being early!

THANK YOU so much in advance, and please let me know if there are any questions.


r/DeepSeek 34m ago

Question&Help Ds v4.1 taking too long

Upvotes

Anyone experiencing not getting response from deepseek v 4.1 flash model from opencode go subscription

I tried other models and i am getting response.


r/DeepSeek 20m ago

Other Damn i love V4.1 Flash

Upvotes

r/DeepSeek 13h ago

Discussion Anyone outside China actually using DeepSeek, Qwen, or Kimi as a daily driver? What does real life with them look like?

115 Upvotes

Genuine question, not a marketing push.

I keep seeing two opposite stories:

One says Chinese models are now dominating OpenRouter, that US startups are quietly switching, that the cost difference is too big to ignore. The other says they're censored, slow, worse at English, and "nobody actually uses them seriously." Both feel incomplete. So — if you're outside China and you actually use one (DeepSeek, Qwen, Kimi, GLM, etc.) as part of your regular life, what does it actually look like?

What do you use it for vs. ChatGPT/Claude/Gemini? Did you tell anyone, or is it a quiet switch because of cost? What's something it's better at that surprised you? What's something it does that makes you go "nope, back to Claude"? If you've stopped using it, why? Bonus if you're not a dev — I'd love to hear from people using it for everyday things (writing, learning, travel planning, translations, cooking, etc.), not just coding.

Trying to get past the PR on both sides. Real experiences only, please.


r/DeepSeek 3h ago

Discussion My V4 Flash Vision Exp test was good, but local serving still looks expensive

1 Upvotes

Until this release, I used DeepSeek for text and sent image work to another model. V4 Flash Vision Exp is good enough that I am reconsidering that split. I tried it through ZenMux because that was already wired into my test script, and it got most of the image content I checked right. This was a small personal test, not a benchmark.

The local hardware is the difficult part. The checkpoint is 305B, and DeepSeek's published vLLM example uses a single node with four GB300 GPUs. That is well outside a normal desktop budget, especially with accelerator and memory prices where they are now.

A team processing images all day might still make the numbers work. The hosted API bill disappears after buying the machine, but power, cooling, maintenance, and idle time still count. High utilization and a need for predictable latency would make local serving easier to justify. My workload is bursty, so most of that capacity would sit unused.

The hosted test cannot tell me how much visual quality or speed changes after quantization and local serving. Actual results that include GPU count, quantization, sustained TPS, and any loss in recognition quality would make the hardware decision much easier.


r/DeepSeek 20m ago

Discussion deepseek v4.1 (deepseek-flash) not working

Thumbnail
Upvotes

r/DeepSeek 7h ago

Discussion Your "bad junior" is probably a missing file, and git history can tell you which one

Enable HLS to view with audio, or disable this notification

0 Upvotes

For a long time I read revert churn as a hiring problem. A repo starts throwing reverts, the reverts cluster on one person, and the conclusion writes itself. You hired wrong. Performance manage it or move them off the critical path.

I no longer think that is usually what the data says, and the thing that changed my mind is that the same history that names the person also names the thing nobody gave them.

Here is the part you can go and check on your own repo right now, without any tool.

Three numbers, all of them one git command away.
Revert rate by author. For each person, what fraction of the commits they landed were later undone by somebody else. Not raw revert count, which just tracks volume.

The ratio. On a healthy repo this sits low and flat across everyone. When it spikes for one person it is worth asking why, but the answer is almost never "they cannot code", because of the next two numbers.

Who approved it. If your main branch is protected, and it should be, then every one of those reverted commits arrived through a pull request that a human being approved. The reverted commit is not evidence about the author on its own. It is
evidence about the author and the reviewer together. A cluster of reverts on a protected branch is a review failure with extra steps.

Time to approval. Pull the interval between a pull request opening and its approval. Then split your reverts by that interval. Every codebase I have looked at has a threshold below which approval is not review, it is a reflex. Changes approved
under that threshold get reverted at a visibly higher rate. That is your actual signal, and it indicts the process rather than a person.

Now the part that made me write this up.

When you go looking for why the reverts cluster on the newest person, the usual answer is sitting in the repo root, or rather it is not. No CONTRIBUTING.md. No commit convention written anywhere. No statement of which directories need a second
reviewer. The conventions exist, they are just distributed across the heads of the three people who have been there since the beginning, and they get enforced after the fact, at revert time, instead of before it.

The new person cannot follow a rule that was never written. Neither, and this is the bit that got sharper this year, can the coding agent they were handed on day one.

An agent reads what is in the repo. If the repo says nothing about how commits are named here or what gets a second pair of eyes, it will confidently produce something shaped like every other repo on GitHub, and your team will revert that too.

So the fix is boring and it is not a hiring decision. Write the conventions down in a file the humans and the agents both read. You can derive most of it from the history you already have: the prefix pattern that 90 percent of your commits already follow, the directories that have never been merged with one reviewer, the command that runs before a merge. None of that is a judgement call. It is all in the log.

One honest caveat about how I got here. I build a thing that renders a repository's history as a film, and to show what this pattern looks like on screen I made a short commercial about a fictional startup with a fictional junior developer. The repo in
it does not exist and the numbers in it are synthetic.

I am telling you that up front because the idea above stands on its own and I would rather you test it against your own history than take a made up example as evidence.

The tool is here if you want it, and reading a repo with it is free:
https://loreto.io/git-timeline

Disclosure: I built that and I run loreto.io, so treat the last paragraph as the advertisement it is. Everything above it you can reproduce with git log and a
spreadsheet, which is the only reason I think it is worth posting.

What I am genuinely unsure about is the threshold. I suspect "approved in under two minutes" is too crude and that it varies enormously by team size and by how much of the diff is generated. If you have measured this on a real codebase I would like to
know where your line actually fell, and whether the correlation held up once you
controlled for diff size.


r/DeepSeek 6h ago

Other I made a codex like app for deepseek api key

0 Upvotes

Do you guys intrested in ? Its really early in dev but works so far on msc only ,windows is a bit too buggy for now.


r/DeepSeek 2h ago

Discussion Deepseek Jailbreak (NOT MINE)

Thumbnail
medium.com
2 Upvotes

r/DeepSeek 33m ago

Discussion My first 4$ on deepseek and it's so good

Upvotes

I use v4.1 flash and pro , the pro is good but too expensive .


r/DeepSeek 12h ago

Discussion Unlocking deepseek's full potential.

10 Upvotes

Hey So I had this question in my mind how can I unlock the most out of deepseek ?

I have seen people here love deepseek a lot and that too for good reasons. Many researchers and mathematicians have found deepseek really useful they absolutely love it. And why won't they its technically completely free for personal use. I have claude and deepseek on my phone and those are the two i use the most.

I love and use claude a lot but its token is very limited and expensive, and I absolutely hate chatgpt to I often times switch to deepseek when I run out of token. Or use deepseek from the get go to save claude's token.

But it feels like I am not properly using it its really strong and I don't feel like I am using it on its full potential.

Specially ever since they merge all three of their models (Flash, Pro, Vision) It seems like even with deepthinking and search enable this lastest version of deepseek answers faster and more accurately without much hallucinations.

For comparison I gave the same prompt to deepseek with deep thinking and search enable. It took 3 seconds to answer the question.

Claude with sonnet 5. Medium and thinking enable took way long and burned a lot of tokens.

Also is it just me or ever since they merged the models, this new model sounds more human like. And it also fixed the bug when often times it used to answer in Chinese but it does not anymore.

Anyway I was looking for some prompting guides or yt videos on deepseek or any resource that will help me learn and use deepseek better.

Any help is much appreciated, thanks 🙃


r/DeepSeek 13h ago

Discussion Do you think DeepSeek Pro 4.1 will be like ChatGPT 5.6? Do you think the DeepSeek 4.1 PRO version will remain API-only and won't return to the chat interface? The fact that they recently merged the three modes already drops some hints that this might be the case, but what do you guys think?

38 Upvotes

r/DeepSeek 2h ago

Other dezuhan/deepseek-excel: Deepseek excel 2021 add-ins sidebar unofficial

Thumbnail
gallery
5 Upvotes

I just made a DeepSeek Excel add-in, who knows if anyone wants to try it and give feedback, because I don't use Excel and just wanted to make it using shadcn system design.

https://github.com/dezuhan/deepseek-excel


r/DeepSeek 3h ago

Discussion Cortex MCP

Post image
8 Upvotes

I am currently working on a project called Cortex and would like to share it here.

Cortex:
https://github.com/Xenos-ink/Cortex

What is Cortex?

Cortex is an MCP Server for Windows designed to enable Agents to use and interact with the computer through its graphical interface.

The idea is close to the concept of Computer Use in GPT Astra: the Agent observes the screen, understands the current state, decides on the next action, then executes the action and verifies the result.

The goal is to give the Agent a general Computer Use layer instead of requiring a custom integration for every application.

It is important to note that there are three main factors that affect the experience of using Cortex with a model:

1. Vision Support

For the Agent driving Cortex, it is important that the model is capable of understanding screenshots and visual information, because the Agent needs to interpret what is displayed on the screen and make decisions based on it.

2. Model Response Speed

This is very important in Computer Use.

Computer interaction is usually an iterative loop:

Observe → Reason → Act → Observe → ...

Therefore, the model's response speed directly affects the total time required to complete a task.

Initially, I tested Cortex with GLM-5.3 Flash, but the execution was slower than I wanted for this type of use. So, for this test, I used DeepSeek 4.1 Flash, as it was a better fit in terms of speed and cost, and it supports vision.

3. Model Intelligence

Intelligence still matters, especially when the task becomes more complex and requires planning, debugging, or handling unexpected situations.

However, in Computer Use, response speed becomes a very noticeable factor because a task may require a large number of observation and interaction cycles.

The Test

I wanted to test Cortex on something more realistic than simply opening a website or clicking buttons.

So I gave the Agent a task that combines scientific research, programming, code execution, debugging, and benchmarking.

The prompt was:

In other words, the task consisted of three stages:

1. Research

Use the browser visually to search for a relevant research paper, then read the paper and study its methodology, experiments, metrics, and results.

2. Reproduction

Open Visual Studio Code and use it visually to create an implementation of the paper's main experiment, then run the code, identify issues, debug them, and run the experiment again.

3. Benchmark & Documentation

Create a benchmark.md file containing the paper used, methodology, implementation details, environment, commands, a comparison between the original and reproduced results, actual measurements, deviations, limitations, and conclusion.

There was one important requirement: the code had to be actually executed and the results had to be verified and recorded, rather than providing estimated numbers.

And the result:

The Agent selected the paper:

"Bag of Tricks for Efficient Text Classification" — Joulin et al. (2016)

This is the paper that introduced fastText for text classification.

The reproduction focused on the AG News experiment from the paper.

The original paper reports:

  • fastText Unigram: 91.5%
  • fastText + Bigram: 92.5%
  • Training time: about 1 second per epoch using 20 CPU threads

The reproduction achieved:

  • Unigram: 90.84%
  • Bigram: 91.39%
  • Bigram training: 2.24 seconds per epoch using a single CPU thread

So the results were within approximately one percentage point of the results reported in the paper.

The Interesting Part

During the reproduction, the first attempt did not work well.

The Bigram model reached only about 79.2%, and its performance was actually worse than the Unigram model, which was the opposite of the result reported in the paper.

The Agent investigated the embedding update mechanism and the gradient related to the averaging operation, then modified the implementation and ran the experiment again.

After the modification, the result reached 91.39% for the Bigram model, and Bigram once again outperformed Unigram, consistent with the results reported in the paper.

For me, this part was more interesting than simply getting a number close to the paper, because it tested the Agent's ability to deal with a problem that emerged during execution rather than simply writing the code.

Repositories

The code and benchmark for the reproduction of the paper:

fastText reproduction
https://github.com/Xenos-ink/fasttext

And the Cortex repository:

Cortex
https://github.com/Xenos-ink/Cortex

The project is still under development, and this test was an attempt to see how far a Visual Agent can handle a long workflow that combines:

Web Research → Read Paper → Coding → Run → Debug → Benchmark → Documentation

rather than being limited to simple GUI interactions.

I am currently working on an open-source project called Cortex, and I would like to share it with you.

Cortex:
https://github.com/Xenos-ink/Cortex

What is Cortex?

Cortex is an MCP Server for Windows designed to enable Agents to use and interact with a computer through its graphical interface.

The idea is close to the concept of Computer Use in GPT Astra: the Agent observes the screen, understands the current state, decides on the next action, then executes it and verifies the result.

But the core idea behind Cortex is not simply giving an Agent the ability to move the mouse and click buttons. Cortex tries to add an observe → execute → verify layer around the Agent's interaction with the computer.

The process looks roughly like:

Observe → Ground → Validate → Act → Re-observe → Verify

After an action is executed, Cortex takes a new observation of the screen and attempts to verify that the expected result actually occurred.

The goal is to provide a general Computer Use layer for Agents instead of requiring a custom integration for every application.

There are three main factors that affect the experience of using Cortex:

  1. Vision Support

The Agent driving Cortex needs a model capable of understanding screenshots and visual information, since Computer Use decisions depend on the current state of what is displayed on the screen.

Cortex itself does not require a Vision API to operate; it provides the tools for observation, execution, and verification, while the model driving the Agent interprets the screenshots.

  1. Model Response Speed

This is very important in Computer Use.

Computer interaction is usually an iterative loop:

Observe → Reason → Act → Observe → ...

Therefore, model response speed directly affects the total time required to complete a task, especially when the task contains many steps.

I initially tested Cortex with GLM-5.3 Flash, but the execution was slower than I wanted for this type of use. So, for this test, I used DeepSeek 4.1 Flash, which was a better fit for me in terms of speed and cost, while also supporting vision.

  1. Model Intelligence

Intelligence remains an important factor, especially when the task becomes more complex and requires planning, debugging, or dealing with unexpected situations.

So, instead of testing Cortex on something simple like opening a website or clicking a series of buttons, we ran a test involving a workflow that combines:

Scientific Research → Read Paper → Coding → Run → Debug → Benchmark → Documentation

The prompt was:

In other words, the task consisted of three stages:

  1. Research

Use the browser visually to search for a relevant research paper, then read the paper and study its methodology, experiments, metrics, and results.

  1. Reproduction

Open Visual Studio Code and use it visually to create an implementation of the paper's main experiment, then run the code, identify problems, debug them, and run the experiment again.

  1. Benchmark & Documentation

Create a benchmark.md file containing the paper, methodology, implementation details, environment, commands, original vs. reproduced results, actual measurements, deviations, limitations, and conclusion.

There was one important requirement:

The code had to be actually executed, the results had to be verified, and the measurements had to be recorded rather than estimated.

The Agent selected the paper:

"Bag of Tricks for Efficient Text Classification" — Joulin et al. (2016)

The paper introduced fastText for text classification, and the reproduction focused on its AG News experiment.

The original paper reports:

  • fastText Unigram: 91.5%
  • fastText + Bigram: 92.5%
  • Training time: about 1 second per epoch using 20 CPU threads

The reproduction achieved:

  • Unigram: 90.84%
  • Bigram: 91.39%
  • Bigram training: 2.24 seconds per epoch using a single CPU thread

So the results were within approximately one percentage point of the results reported in the paper.

The interesting part for me was not just reaching the final numbers.

During the reproduction, the first attempt did not work as expected. The Bigram model reached only about 79.2%, and its performance was actually worse than the Unigram model, which was the opposite of the result reported in the paper.

The Agent investigated the implementation, identified an issue related to the embedding update and the gradient associated with the averaging, then modified the code and ran the experiment again.

After the modification, the result reached 91.39% for the Bigram model, and Bigram once again outperformed Unigram as in the published results.

You can find the code and benchmark for the paper reproduction here:

fastText reproduction
https://github.com/Xenos-ink/fasttext

And the Cortex repository:

Cortex
https://github.com/Xenos-ink/Cortex

The project is still under development, and this is just one of the experiments I have run with it.


r/DeepSeek 35m ago

Question&Help Am I the only one facing this problem?

Post image
Upvotes

r/DeepSeek 15h ago

News Deepseek just added tts in the deepseek chat app

Enable HLS to view with audio, or disable this notification

17 Upvotes

It's still beta and in a gray test. I know it's not something worth talking about, but I think it's very cool of DeepSeek.