Showing posts with label DeepSeek. Show all posts
Showing posts with label DeepSeek. Show all posts

Saturday, February 15, 2025

DeepSeek explained: Everything you need to Know

In the world of AI, there has been a prevailing notion that developing leading-edge large language models requires significant technical and financial resources. That's one of the main reasons why the U.S. government pledged to support the $500 billion Stargate Project announced by President Donald Trump.

 

But Chinese AI development firm DeepSeek has disrupted that notion. On Jan. 20, 2025, DeepSeek released its R1 LLM at a fraction of the cost that other vendors incurred in their own developments. DeepSeek is also providing its R1 models under an open source license, enabling free use.

Within days of its release, the DeepSeek AI assistant -- a mobile app that provides a chatbot interface for DeepSeek-R1 -- hit the top of Apple's App Store chart, outranking OpenAI's ChatGPT mobile app. The meteoric rise of DeepSeek in terms of usage and popularity triggered a stock market sell-off on Jan. 27, 2025, as investors cast doubt on the value of large AI vendors based in the U.S., including Nvidia. Microsoft, Meta Platforms, Oracle, Broadcom and other tech giants also saw significant drops as investors reassessed AI valuations.

 

What is DeepSeek?

DeepSeek

DeepSeek is an AI development firm based in Hangzhou, China. The company was founded by Liang Wenfeng, a graduate of Zhejiang University, in May 2023. Wenfeng also co-founded High-Flyer, a China-based quantitative hedge fund that owns DeepSeek. Currently, DeepSeek operates as an independent AI research lab under the umbrella of High-Flyer. The full amount of funding and the valuation of DeepSeek have not been publicly disclosed.

 

DeepSeek focuses on developing open source LLMs. The company's first model was released in November 2023. The company has iterated multiple times on its core LLM and has built out several different variations. However, it wasn't until January 2025 after the release of its R1 reasoning model that the company became globally famous.

The company provides multiple services for its models, including a web interface, mobile application and API access.

OpenAI vs. DeepSeek

DeepSeek Vs OpenAI


DeepSeek represents the latest challenge to OpenAI, which established itself as an industry leader with the debut of ChatGPT in 2022. OpenAI has helped push the generative AI industry forward with its GPT family of models, as well as its o1 class of reasoning models.

 

While the two companies are both developing generative AI LLMs, they have different approaches.

OpenAIDeepSeek
Founding year20152023
HeadquartersSan Francisco, Calif.Hangzhou, China
Development focusBroad AI capabilitiesEfficient,
open source models
Key modelsGPT-4o, o1DeepSeek-V3, DeepSeek-R1
Specialized modelsDall-E (image generation),
Whisper (speech recognition)
DeepSeek Coder (coding), Janus Pro (vision model)
API pricing
(per million tokens)
o1: $15 (input), $60 (output)DeepSeek-R1: $0.55 (input), $2.19 (output)
Open source policyLimitedMostly open source
Training approachSupervised and instruction-based fine-tuningReinforcement learning
Development costHundreds of millions of dollars for o1 (estimated)

Less than $6 million for DeepSeek-R1, according to the company

Training innovations in DeepSeek

DeepSeek uses a different approach to train its R1 models than what is used by OpenAI. The training involved less time, fewer AI accelerators and less cost to develop. DeepSeek's aim is to achieve artificial general intelligence, and the company's advancements in reasoning capabilities represent significant progress in AI development.

 

In a research paper, DeepSeek outlines the multiple innovations it developed as part of the R1 model, including the following:

  • Reinforcement learning. DeepSeek used a large-scale reinforcement learning approach focused on reasoning tasks.
  • Reward engineering. Researchers developed a rule-based reward system for the model that outperforms neural reward models that are more commonly used. Reward engineering is the process of designing the incentive system that guides an AI model's learning during training.
  • Distillation. Using efficient knowledge transfer techniques, DeepSeek researchers successfully compressed capabilities into models as small as 1.5 billion parameters.
  • Emergent behavior network. DeepSeek's emergent behavior innovation is the discovery that complex reasoning patterns can develop naturally through reinforcement learning without explicitly programming them.

 

DeepSeek large language models
China Deepseek
Since the company was created in 2023, DeepSeek has released a series of generative AI models. With each new generation, the company has worked to advance both the capabilities and performance of its models:
  • DeepSeek Coder. Released in November 2023, this is the company's first open source model designed specifically for coding-related tasks.
  • DeepSeek LLM. Released in December 2023, this is the first version of the company's general-purpose model.
  • DeepSeek-V2. Released in May 2024, this is the second version of the company's LLM, focusing on strong performance and lower training costs.
  • DeepSeek-Coder-V2. Released in July 2024, this is a 236 billion-parameter model offering a context window of 128,000 tokens, designed for complex coding challenges.
  • DeepSeek-V3. Released in December 2024, DeepSeek-V3 uses a mixture-of-experts architecture, capable of handling a range of tasks. The model has 671 billion parameters with a context length of 128,000.
  • DeepSeek-R1. Released in January 2025, this model is based on DeepSeek-V3 and is focused on advanced reasoning tasks directly competing with OpenAI's o1 model in performance, while maintaining a significantly lower cost structure. Like DeepSeek-V3, the model has 671 billion parameters with a context length of 128,000.
  • Janus-Pro-7B. Released in January 2025, Janus-Pro-7B is a vision model that can understand and generate images.

 

Why it is raising alarms in the U.S

While there was much hype around the DeepSeek-R1 release, it has raised alarms in the U.S., triggering concerns and a stock market sell-off in tech stocks. On Monday, Jan. 27, 2025, the Nasdaq Composite dropped by 3.4% at market opening, with Nvidia declining by 17% and losing approximately $600 billion in market capitalization.

 

DeepSeek is raising alarms in the U.S. for several reasons, including the following:

  • Cost disruption. DeepSeek claims to have developed its R1 model for less than $6 million. The low-cost development threatens the business model of U.S. tech companies that have invested billions in AI. DeepSeek is also cheaper for users than OpenAI.
  • Technical achievement despite restrictions. The export of the highest-performance AI accelerator and GPU chips from the U.S. is restricted to China. Yet, despite that, DeepSeek has demonstrated that leading-edge AI development is possible without access to the most advanced U.S. technology.
  • Business model threat. In contrast with OpenAI, which is proprietary technology, DeepSeek is open source and free, challenging the revenue model of U.S. companies charging monthly fees for AI services.
  • Geopolitical concerns. Being based in China, DeepSeek challenges U.S. technological dominance in AI. Tech investor Marc Andreessen called it AI's "Sputnik moment," comparing it to the Soviet Union's space race breakthrough in the 1950s.

 

DeepSeek Bans
DeepSeek Ban
Countries and organizations around the world have already banned DeepSeek, citing ethics, privacy and security issues within the company. Because all user data is stored in China, the biggest concern is the potential for a data leak to the Chinese government. The LLM was also trained with a Chinese worldview -- a potential problem due to the country's authoritarian government.

Places where DeepSeek is banned include the following:

  • Australian government agencies.
  • India central government.
  • Italy.
  • NASA.
  • South Korea industry ministry.
  • Taiwan government agencies.
  • Texas state government.
  • U.S. Congress.
  • U.S. Navy.
  • U.S. Pentagon.

 

DeepSeek Cyberattack


DeepSeek's popularity has not gone unnoticed by cyberattackers.

On Jan. 27, 2025, DeepSeek reported large-scale malicious attacks on its services, forcing the company to temporarily limit new user registrations. The timing of the attack coincided with DeepSeek's AI assistant app overtaking ChatGPT as the top downloaded app on the Apple App Store.

Despite the attack, DeepSeek maintained service for existing users. The issue extended into Jan. 28, when the company reported it had identified the issue and deployed a fix.

DeepSeek has not specified the exact nature of the attack, though widespread speculation from public reports indicated it was some form of DDoS attack targeting its API and web chat platform.

 

DeepSeek data Exposed

Wiz Research -- a team within cloud security vendor Wiz Inc. -- published findings on Jan. 29, 2025, about a publicly accessible back-end database spilling sensitive information onto the web -- a "rookie" cybersecurity mistake. Information included DeepSeek chat history, back-end data, log streams, API keys and operational details. DeepSeek took the database offline shortly after being informed. It's unclear for how long the database was exposed.

 

 

 

Wednesday, February 12, 2025

How to Run DeepSeek Locally on Your Machine: A Step-by-Step Guide

How to Run DeepSeek Locally on Your Machine: A Step-by-Step Guide

How to Run DeepSeek Locally

In the ever-evolving world of artificial intelligence and machine learning, running powerful AI models locally on your machine has become increasingly accessible. DeepSeek, a cutting-edge AI framework, is no exception. Whether you're a developer, data scientist, or AI enthusiast, running DeepSeek locally can provide you with unparalleled flexibility and control over your AI projects. In this blog, we'll walk you through the process of setting up and running DeepSeek on your local machine, ensuring you can harness its full potential.

 

Why Run DeepSeek Locally?

Before diving into the technical details, let's explore why running DeepSeek locally is beneficial:

  1. Privacy and Security: By running DeepSeek locally, you ensure that your data remains on your machine, reducing the risk of data breaches.
  2. Customization: Local execution allows you to tweak and customize the model to suit your specific needs.
  3. Offline Access: Once set up, you can use DeepSeek without an internet connection, making it ideal for environments with limited connectivity.
  4. Performance: Running the model locally can often result in faster processing times, especially if you have a powerful machine.

Prerequisites

 

Before we begin, ensure that your machine meets the following requirements:

  1. Operating System: Windows, macOS, or Linux

  2. Python: Version 3.7 or higher

  3. GPU: Optional but recommended for faster processing (NVIDIA GPU with CUDA support)

  4. RAM: At least 8GB (16GB or more recommended)

  5. Storage: Sufficient space for the model and datasets (SSD recommended for faster access)

Step 1: Install Python and Required Libraries

 

First, ensure that Python is installed on your machine. You can download the latest version of Python from the  official Python website.

Once Python is installed, open your terminal or command prompt and install the necessary libraries using pip:

bash

pip install torch torchvision torchaudio
pip install transformers
pip install deepseek


These libraries include PyTorch, which is essential for running DeepSeek, and the transformers library, which provides pre-trained models and utilities for natural language processing.

 

Step 2: Download the DeepSeek Model

Next, you'll need to download the DeepSeek model. You can do this using the transformers library:


python

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepseek/deepseek-llm"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)


This code snippet downloads the DeepSeek model and its corresponding tokenizer, which are essential for processing text inputs.

 

Step 3: Set Up Your Environment

To ensure optimal performance, it's crucial to set up your environment correctly. If you have an NVIDIA GPU, make sure that CUDA and cuDNN are installed. You can verify this by running:

bash

nvidia-smi


If your GPU is recognized, you can enable GPU acceleration by moving the model to the GPU:

python

import torch

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

Step 4: Run DeepSeek Locally

 

Now that everything is set up, you can run DeepSeek locally. Here's a simple example of how to generate text using the DeepSeek model:

python

input_text = "Once upon a time"
inputs = tokenizer(input_text, return_tensors="pt").to(device)
outputs = model.generate(**inputs, max_length=50)

generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(generated_text)


 

This code takes an input text, processes it using the DeepSeek model, and generates a continuation of the text. The max_length parameter controls the length of the generated text.

Step 5: Optimize and Fine-Tune

 

Running DeepSeek locally allows you to fine-tune the model on your specific dataset. Fine-tuning can significantly improve the model's performance on your particular use case. Here's a basic example of how to fine-tune the model:


python

from transformers import Trainer, TrainingArguments

training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    per_device_eval_batch_size=4,
    warmup_steps=500,
    weight_decay=0.01,
    logging_dir="./logs",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=your_train_dataset,
    eval_dataset=your_eval_dataset,
)

trainer.train()


Replace your_train_dataset and your_eval_dataset with your actual datasets. Fine-tuning can take some time, especially if you're working with a large dataset, but the results are often worth the effort.

 

Conclusion

Running DeepSeek locally on your machine is a powerful way to leverage AI for your projects. By following the steps outlined in this guide, you can set up, run, and even fine-tune the DeepSeek model to meet your specific needs. Whether you're developing AI applications, conducting research, or simply exploring the capabilities of AI, running DeepSeek locally offers a level of control and flexibility that cloud-based solutions simply can't match.

 

So, what are you waiting for? Dive into the world of DeepSeek and unlock the full potential of AI on your local machine today!

 

 

 



Monday, February 3, 2025

Best AI Assistants (Chatbots)

The Best AI assistants (Chatbots)

AI Tools
AI Assistants (Chatbots): 
Ch‎atGPT
ChatGPT

ChatGPT consistently ranks at the top of the LM Arena leaderboard, outperforming other models in key benchmarks. It's the world's most popular AI application, with 200 million users as of October 2024.

I’ve used ChatGPT extensively for brainstorming ideas, translation tasks, coding, AI script generation, data analysis and managing research-heavy tasks. The new 4o model is a significant leap forward—it’s incredibly fast, and feels way smarter than any of the previous versions of ChatGPT.

With ChatGPT's multimodal capabilities I can paste in images—like a chart or graph—and ask questions about it, making it much easier to interpret visual data quickly. I fed it a PNG image of a chart and it analysed the chart, gave me a table of the raw data (that it read from the chart image) and then re-did the chart in my preferred colors - pretty impressive.

ChatGPT can now retain context over time, personalizing responses based on previous conversations. For instance, I’ve used it to refine recurring project ideas without re-explaining every detail, saving hours of effort. You can review and manage what it remembers through OpenAI’s controls, to make sure it doesn't go all Skynet on you.

The integrated ChatGPT search option (more on this later) makes it even easier to find relevant information directly within conversations, which cuts down on the hallucinations with the use of RAG (Retrieval Augmented Generation). RAG grounds the AI's answer by retrieving information from external data sources.

While it excels in creative and general-purpose tasks, I’d recommend exploring other tools like Claude (see below) for coding. It's not that ChatGPT is bad at coding tasks, it's just that Claude is great at them.

ChatGPT o1

o1 is a specialized advanced reasoning model built for complex problem-solving, coding, and math.

While I find 4o excels in creativity and versatility, o1 has proven incredibly useful for specific tasks like coding, troubleshooting technical issues, and even solving intricate math problems that other models struggled with. I’ve used it to generate shell scripts, work through spreadsheet problems, and even tackle cryptic crossword puzzles, where its precision and logical depth really shine.

However, it lacks the broader capabilities and tool integrations of 4o, so I see it more as a complementary option for specific needs rather than a full replacement for more creative or expansive tasks.

Operators

In January 2025, ChatGPT introduced "Operators," AI agents that can book hotels, order food, and shop online. Exclusive to Pro users ($200/month), they show exciting potential but are hit-or-miss in execution.

For instance, I asked the Operator to book a hotel in NYC. It started strong, navigating filters and searching TripAdvisor, but eventually got stuck in a loop. Ordering a pizza was similar—it customized the order but couldn’t complete checkout. Shopping worked better; it found a laptop under $1,000 on Amazon but required me to finish the purchase manually. Operators let you take control when they get stuck, but the laggy browser often makes doing it yourself easier.

Right now, Operators feel more like a proof-of-concept than a practical tool. While the idea of automating repetitive tasks is exciting, it needs improvements in speed and reliability. If you’re already a Pro user, it’s worth exploring, but not essential yet.

Pricing

OpenAI offer a free tier which currently gives you limited access to GPT-4o and unlimited access to ChatGPT-4o mini. The Plus plan gets you wider access and costs $20/month - I think that's pretty good value for money. They also offer a Pro plan for $200/month which gives you priority access to their latest tools.

Cl‎aude

Claude

I’ve been using Claude (their Sonnet 3.5 model to be specific), for coding tasks, and it’s quickly becoming my go-to for code reviews. What really makes Claude stand out is how precise it is—it seems to "get" the nuances of programming better than other tools I’ve tried. I’ve used it to spot subtle issues in my code and even brainstorm better ways to structure projects. Anthropic are training these models on more recent and specialized coding knowledge and it shows, especially when tackling modern frameworks or troubleshooting tricky bugs.

Another thing I love about Claude is how nice it is to talk to. It feels like it has more "soul" compared to ChatGPT—the tone is warmer, and conversations just flow better. Whether I’m bouncing around ideas or working through a complicated issue, it’s genuinely pleasant to interact with. I have quite reached Her levels of affection for Claude, but we're getting there.

That said, I have hit the response and rate limits a little faster than I’d like, which can be a hassle if I’m deep into a project. But for $20/month on the Pro plan, it’s still a great deal, especially if you’re looking for an AI assistant that’s smart, approachable, and particularly strong in coding tasks.

Gemini

Gemini

Google’s Gemini fits seamlessly into the Google ecosystem. On Android, it feels like a natural extension of the system rather than a separate app, and if you’re already using Google Workspace, it’s incredibly convenient. Whether I was drafting emails, summarizing articles, or asking it random questions, it delivered quickly and smoothly.

I’ve found it useful in unexpected ways too. When reviewing legal documents, I’d do my initial read-through and then ask Gemini to double-check if I missed anything. Another time, I struggled with a confusing sizing chart while shopping for clothes. I snapped a picture of the label, described my usual size, and let Gemini handle the rest. The suggestion was spot-on, and I ended up with a perfect fit!

For creative projects, Gemini’s image capabilities really shine. I once uploaded an image I liked and asked it to describe it as a prompt for an AI image generator. The results were creative and inspiring, making it a fun tool for brainstorming new ideas.

When I was working on a project proposal, Gemini Advanced provided nuanced, tailored suggestions that felt like a genuine productivity boost. It even made copywriting easier—generating meaningful text for design mockups that felt polished, rather than using generic filler like "Lorem Ipsum."

However, it’s not perfect. One frustration I had was with its context retention. When revising a piece of writing, I had to re-explain instructions a few times because it would forget what we’d already discussed. Similarly, when I uploaded an Excel file, got a summary, and later updated the data, Gemini treated the updated file as a brand-new task instead of building on what we’d already done.

Another weak spot is its performance on technical tasks. While it’s great at formatting and debugging simple code, I found that it sometimes rewrote JavaScript as Python unnecessarily. For more specialized or dense content, like legal texts, its analysis lacked depth compared to what I was hoping for. Even its responses to some image-based queries were occasionally inaccurate, which was a letdown after seeing its creative potential elsewhere.

That said, Gemini’s strengths outweigh its flaws. Its tight integration with Google tools makes it practical for anyone already in the Google ecosystem, and its ability to handle both text and images makes it a versatile tool for creative projects. While it’s not the best choice for highly technical or niche tasks, it’s a solid, fast, and easy-to-use assistant for everyday needs—and for me, the advanced features have made it a tool I’ve come to rely on.

While the free Basic version (using the 1.5 Flash model) covers most casual needs, the $19.99/month Gemini Advanced adds the more powerful 1.5 Pro and Gemini-Exp-1206 models for complex tasks like coding, math, and deep research, including analyzing texts up to 1,500 pages.

De‎epSeek

DeepSeek

DeepSeek is also worth checking out. They let you use their V3 and new R1 models for free on their site, although you still have to pay for API access (it's very cheap though).

DeepSeek's search feels more engaging and "sticky" even after just a few queries. Its transparency—showing reasoning and openly acknowledging what it knows and what it might not—builds a significant level of user trust.

In January 2025, they launched their R1 model as a competitor to ChatGPT's o1, quickly gaining attention in the AI community for being both cost-effective and open source. I've played around with both their R1 and V3 models.

I asked both ChatGPT-o1 and DeepSeek-R1 to analyze sections of a presentation I’m working on. R1 provided a more comprehensive analysis, addressing key aspects that o1 overlooked. I also had both brainstorm ideas, and once again, R1 delivered significantly better suggestions than o1.

For coding I’ve been relying more on DeepSeek (v3) lately because of its straightforward approach—it gets straight to the point with its suggestions. Claude (3.5 Sonnet), by contrast, often takes a more detailed route, proposing multiple solutions and leaning toward the one that aligns best with solid software engineering practices. Both tools are excellent in their own ways, and I’ve started using them equally. DeepSeek is great for its affordability and efficiency, while Claude is invaluable for double-checking critical code and ensuring everything is on point. Together, they make a great team.

For writing, I'm less keen on these DeepSeek models. I find its output less natural-sounding and oftentimes boring and repetitive.