Researchers Suspect DeepSeek Cloned Gemini Data

Observers noted that DeepSeek R1-0528 mimics Gemini’s phrasing and reasoning style.

DeepSeek, a Chinese lab, released an updated version of its R1 reasoning AI model last week. The company did not disclose the data sources used for training, but some AI researchers suggest that Google’s Gemini family of AI may have been a source.

Sam Paech, a Melbourne-based developer, claims to have found evidence that DeepSeek’s latest model was trained on outputs from Gemini. According to Paech’s X post, DeepSeek’s model, R1-0528, uses similar words and expressions favored by Google’s Gemini 2.5 Pro.

SpeechMap’s pseudonymous creator, who developed a “free speech eval” for AI, mentioned that the DeepSeek model’s “thoughts” resemble Gemini traces. Previously, DeepSeek faced accusations of training on data from competitor AI models. In December, developers noticed that DeepSeek’s V3 model often identified itself as ChatGPT, suggesting potential training on ChatGPT chat logs.

Earlier in the year, OpenAI informed the Financial Times about evidence connecting DeepSeek to distillation, a technique involving extracting data from larger AI models for training. According to Bloomberg, Microsoft detected significant data exfiltration through OpenAI developer accounts in late 2024, accounts OpenAI suspects are linked to DeepSeek.

OpenAI prohibits customers from using its model outputs to create competing AI, despite the fact that distillation is relatively common. AI companies source training data from the open web, increasingly saturated with AI-generated content. This has made it difficult to thoroughly filter AI outputs from training datasets.

Nathan Lambert, a researcher at AI2, believes that DeepSeek may have trained on data from Google’s Gemini. Lambert stated in an X post, “If I was DeepSeek, I would definitely create a ton of synthetic data from the best API model out there… [DeepSeek is] short on GPUs and flush with cash. It’s literally effectively more compute for them.”

AI companies are increasing security measures to prevent distillation. OpenAI began requiring organizations to complete an ID verification process to access advanced models in April. China is not on the list of countries supported by OpenAI’s API for this process.

Google has begun “summarizing” traces generated by models available through its AI Studio developer platform. Anthropic announced plans in May to summarize its own model’s traces.

Featured image credit

Researchers suspect DeepSeek cloned Gemini data

Observers noted that DeepSeek R1-0528 mimics Gemini’s phrasing and reasoning style.

Related Posts

Xbox tests free ad-supported cloud gaming

OpenAI launches ChatGPT Health to all US users

Runway introduces AI model router via Dev platform

AMD unveils Helios AI rack to challenge Nvidia

Amazon brings Luna games into Prime Video

Anthropic upgrades Claude voice mode with Sonnet

LATEST NEWS

Xbox tests free ad-supported cloud gaming

OpenAI launches ChatGPT Health to all US users

Runway introduces AI model router via Dev platform

AMD unveils Helios AI rack to challenge Nvidia

Amazon brings Luna games into Prime Video

Anthropic upgrades Claude voice mode with Sonnet

BEST AI MODELS LEADERBOARD

LATEST TOOLS

Amanda AI

InterviewBot

VernAI

MyLoans

Essay Grader AI

Cover Letter AI

Animate Old Photos

Resume.io

MonAI

AIEngine Plugin

COPYRIGHT © DATACONOMY MEDIA GMBH, ALL RIGHTS RESERVED.

Researchers suspect DeepSeek cloned Gemini data

Observers noted that DeepSeek R1-0528 mimics Gemini’s phrasing and reasoning style.

Stay Ahead of the Curve!

Related Posts

LATEST NEWS

COPYRIGHT © DATACONOMY MEDIA GMBH, ALL RIGHTS RESERVED.

Follow Us