[Photo: Reve AI]

[Digital Today reporter Chi-gyu Hwang] Competition among tech companies is heating up over lightweight models, following the development of ultra-large models with nearly 10 trillion parameters.

As AI-related costs rise, more companies are deploying cost-effective open-source models for basic tasks. In the process, local models that can run on personal PCs appear to be gaining prominence.

Moves by major companies targeting local models are also accelerating. In China, DeepSeek, and in the United States, Meta have recently unveiled models that can be installed and used on PCs while still offering solid performance. Nvidia is also speeding up efforts to expand its lineup of AI models for local environments.

DeepSeek released V4 Flash, a new model aimed at coding with 304 billion parameters, in late July. V4 Flash is smaller than DeepSeek's flagship V4 Pro (DeepSeek-V4-Pro) preview, which has about 1.6 trillion parameters, but it is drawing interest among developers for offering performance comparable to large models.

According to reports, V4 Flash showed performance close to Claude Opus 4.8 and far ahead of the V4 Pro preview version even though it went through minimal fine-tuning. In tests for complex coding and autonomous software tasks, it received results comparable to Anthropic's Claude Opus 4.8, and it outperformed Opus 4.8 on the Arena frontend coding leaderboard.

Considering pricing, V4 Flash is viewed as even more powerful. Opus 4.8 costs $25 per minute of output, while the price of the V4 Flash API offered by DeepSeek is 28 cents. That is 99 percent cheaper.

Developers have also been talking about V4 Flash being able to run on consumer hardware. Advanced AI models require servers with hundreds of gigabytes of GPU memory, but reviews say V4 Flash ran well enough on high-end personal PCs. Salvatore Sanfilippo, a developer known by the pseudonym antirez and the creator of the Redis database system, said he is considering switching to V4 Flash from GLM 5.2 for small computers.

Meta has presented a more mainstream local AI strategy. Meta also plans to release a new model family, Muse Glimmer, that can run on laptops.

Muse Glimmer is a 30 billion-parameter model, and its weights will be distributed under the Apache 2.0 licence so developers can download and modify them. Muse Glimmer is designed for AI agents that carry out multi-step tasks. Using a single consumer GPU on a Mac or PC, it can handle tool calling, code writing and debugging, file and screenshot tasks, and long-running workflows. It supports text and images and was trained in more than 100 languages.

Meta has emphasised privacy in connection with Muse Glimmer's target market. Meta's plan is to use Muse Glimmer to process information on users' devices without sending it to the cloud, laying the groundwork for personal agent services that are sensitive to privacy. It presented tasks such as schedule management, drafting messages and organising files, which require access to personal data, as examples of how Muse Glimmer can be used.

Meta CEO Mark Zuckerberg (마크 저커버그) said, "Personal agents will help users 24 hours a day with relationships, health, career, finances, household management and hobbies, and everyone will have access to these tools for free or at an affordable cost."

There is still a barrier to using Muse Glimmer on PCs. Muse Glimmer requires a GPU with at least 24GB of VRAM, which could make large-scale adoption difficult for companies. Computerworld reported that analysts and consultants agree companies show strong interest in local operation, but assessing whether a switch from cloud to local is financially justified is far more complex, adding that a key issue is that it is impossible to predict how RAM prices and cloud prices will move over the next 12 to 18 months.

Nvidia also unveiled the open-source AI model Nemotron 3.5 Lightning, which companies can download, use and modify without permission or fees. Nemotron 3.5 Lightning is lightweight and can run on PC GPUs.

Nemotron 3.5 Lightning is a 30 billion-parameter mixture-of-experts model developed for agents, AI programs that operate autonomously in the background, and it is available on Hugging Face and Nvidia's website. Companies can use Nvidia NeMo on their own hardware to conduct follow-on training tailored to in-house domain data, tools and workflows.

Keyword

#DeepSeek #Meta #Nvidia #Muse Glimmer #Nemotron 3.5 Lightning
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.