Vercel AI SDK 5 Ships Voice API; Ollama Gets a GUI and Multimodal Chops
Hello, developer friends:
This is the 「RTE Developer Daily」 , bringing you news and gossip every day. Our community editorial team curates and shares "topical technologies", "highlight products", "thoughtful articles", "opinionated viewpoints", and "noteworthy events" in the RTE (Real-Time Engagement) field. However, the content only represents the editor's personal views. Everyone is welcome to leave comments, follow threads, and discuss.
Editors for this issue: @赵怡岭, @鲍勃
01 Topical Technologies
1. Veo 3 Fast is coming to Google AI Studio, supporting the generation of dynamic videos with desired actions, narratives, and sound effects
On July 30, Google announced a new addition to its video generation model family, Veo 3 Fast, and added image-to-video generation capabilities for Veo 3 and Veo 3 Fast. These updates are now available as a paid preview through the Gemini API.
Veo 3 Fast is an optimized version of the Veo 3 model, designed for developers pursuing speed and cost-effectiveness, enabling faster creative iteration. The model supports both text-to-video and image-to-video modes and is well-suited for scenarios like programmatic advertising, rapid creative prototyping, and large-scale social media content generation. Veo 3 Fast is priced at $0.40 per second of video (including audio).
The new image-to-video feature allows developers to use Veo 3 and Veo 3 Fast to generate high-quality video clips with sound from static input images.
Veo 3 can generate video and audio in one step. Users only need to provide an image and a corresponding text prompt to guide the model in generating a dynamic video with desired actions, narratives, and sound effects, while maintaining stylistic consistency with the initial image. This feature is priced the same as text-to-video on Veo 3, at $0.75 per second of video (including audio).
Additionally, there are reports that Veo 3 Fast may soon also land on Google AI Studio with a quota.
Related link: https://developers.googleblog.com/en/veo-3-fast-image-to-video-capabilities-now-available-gemini-api/ (@橘鸭 Juya, @WebEye Cloud Services)
2. Vercel releases AI SDK 5, introducing a Voice API and enhanced tool calling
Vercel has released the fifth major version of its popular open-source AI application toolkit, AI SDK. With over 2 million weekly downloads, AI SDK 5 brings several major updates to TypeScript and JavaScript, including type-safe chat functionality, Agentic loop control, voice generation and transcription, and enhanced tooling.
The new version has completely rebuilt the chat functionality, with the core introduction of two distinct message types to solve developer challenges in state management and chat history persistence. UIMessage serves as the "source of truth" for application state, containing all messages, metadata, and tool results, and is recommended for persistent storage.
The new version also extends the unified provider abstraction to the voice domain, offering a unified, type-safe interface for voice generation and transcription services from providers like OpenAI, ElevenLabs, and DeepGram. Tool calling has also been comprehensively enhanced, supporting dynamic tools, provider-executed tools (like OpenAI's web search), and more granular lifecycle hooks.
Other important updates include: adopting SSE as the standard streaming protocol; introducing a global provider (defaulting to Vercel AI Gateway) to simplify model ID usage; supporting access to raw request and response data for enhanced debugging and control; and adding support for Zod 4. To help users migrate smoothly, Vercel provides an automated code modification tool (@codemods).
Related link: https://vercel.com/blog/ai-sdk-5 (@aisdk on X, @橘鸭 Juya)
02 Highlight Products
1. Manus launches "Wide Research"
On the evening of July 29, Manus officially launched a new feature capable of coordinating dozens of AI agents to work simultaneously for broad research, named "Wide Research".
According to the official introduction, the key to Wide Research is not just having more agents—but how they collaborate. Unlike traditional multi-agent systems based on predefined roles (like "manager", "programmer", "designer"), each sub-agent in Wide Research is a fully functional, general-purpose Manus instance.
Reportedly, with Wide Research, Manus unlocks a powerful new way for users to handle complex, large-scale tasks that require gathering information on hundreds of items. "Whether you are exploring Fortune 500 companies, comparing top MBA programs, or delving into GenAI tools, Wide Research makes deep, extensive research effortless."
Wide Research is officially available to Pro users starting today, with plans to gradually roll out to Plus and Basic tier users.
Experience address: https://manus.im/app (@APPSO)
2. Moonvalley launches Sketch-to-Video feature: hand-drawn sketches instantly generate cinematic videos
AI video generation startup Moonvalley recently announced that its flagship model, Marey, officially supports a Sketch-to-Video feature. This means users can simply draw a sketch to have the system quickly generate dynamic footage with a cinematic feel.
This feature is an important extension of Marey's "hybrid creation" concept. Unlike traditional generation methods that only accept text prompts, users can quickly define scene structure, poses, or composition through hand-drawn frames, which the Marey model then converts into specific shots. This experience aligns more closely with a director's visual creation workflow, helping creators bring their ideas to life faster.
Moonvalley emphasizes that this feature provides professional creators with more interactive control. For example, after defining character movements or camera motion paths through a sketch, the system can automatically generate coherent video clips, supporting detailed post-production adjustments. The model supports 1080p@24fps output, ensuring picture clarity and smoothness.
This step marks AI video production's shift from "black-box generation" towards "director-like control," making Marey a professional tool more aligned with film and television workflows. Moonvalley CEO Naeem Talukdar stated: "We want the tool to organically integrate with the director's creative process, not simply replace it."
Currently, the Sketch-to-Video feature is available to subscribers through the Marey platform, with subscription prices starting at $14.99/month; users can also choose to purchase rendering credits on demand. Moonvalley is collaborating with several film and television institutions and advertising teams to explore its application value in commercial production workflows.
Related link: https://techcrunch.com/2025/07/08/moonvalleys-ethical-ai-video-model-for-filmmakers-is-now-publicly-available/ (@AI Planet Vision)
3. Ollama version 0.10.1 officially launches a visual graphical interface, supporting model downloads and multimodal interaction
Ollama version 0.10.1 officially launches a visual graphical interface, supporting both Mac and Windows simultaneously.
Feature introduction:
Simpler conversation interface: The new version of Ollama provides a brand new conversation interface, supporting not only regular conversations but also model downloads;
Chat with files: The new version of Ollama also supports chatting with PDFs and documents; for larger documents, you can increase Ollama's context length in settings, which may consume more computer memory;
Supports multimodal conversation: The new version of Ollama has a built-in new multimodal engine, supporting sending images to large language models, provided the model supports multimodality, such as Gemma 3, etc., and domestic models like Qwen2.5vl, etc.;
Document writing: The new version of Ollama supports adding code files and then having the large language model understand them and write new documents.
Ollama official website: https://ollama.com/
Download address: https://ollama.com/download (@AI Tool Pie)
4. Generative media platform fal announces $125 million Series C funding led by Meritech Capital
Today, fal announced the completion of a $125 million Series C funding round led by Meritech Capital, with Salesforce Ventures, Shopify Ventures, and Google AI Futures Fund also joining this round. Notably, the company successfully completed three funding rounds within 12 months, and its current valuation reaches $1.5 billion.
fal is a San Francisco-based company dedicated to building the world's first generative media platform for developers.
fal's Generative Media Cloud now supports tens of thousands of applications, serving over two million developers and more than three hundred enterprise customers.
In the past year alone, the company's average monthly growth rate reached 40%. With this latest round of funding, fal plans to significantly expand its engineering, support, sales, and marketing teams to keep up with the growing demands and enthusiasm of its community users.
fal has put forward its development vision: to build a generative media platform that easily creates dynamic, real-time content covering video, audio, images, and 3D.
Related link: https://fal.ai/careers (@FAL on X)
5. AI voice input app Willow now supports transcription that mimics user style
On August 1, the AI voice input app Willow updated its personalization feature to support transcription that mimics user style.
Through the personalization feature, with just a small amount of editing by the user, Willow can mimic the user's editing style, language, tone, etc., and transcribe in the user's personalized style.
Last month, Willow just completed a $4.2 million fundraising round.
Experience link: willowvoice.com (@WillowVoiceAI on X)
6. Coze adds a new agent publishing channel, supporting one-click publishing of agents to the Xiaomi App Store
Starting August 1, the Coze development platform is officially connected with the Xiaomi App Store, adding a new publishing channel—enabling one-click publishing of agents to the Xiaomi App Store, accelerating the landing and dissemination of intelligent creativity.
After passing the Xiaomi App Store review, the agent will be listed in the Xiaomi App Store. It can be found by users through the Xiaomi App Store search bar, or viewed in the Xiaomi App Store's [AI Agent Zone], and the agent service can be invoked instantly without downloading or jumping between apps. Hundreds of millions of Xiaomi terminals instantly become your new traffic pool.
Related link: https://www.coze.cn/home (@Coze)
03 Real-Time AI Demo
1. Odyssey: Machine Dream Video Model
From Odyssey founder @olivercameron on X: What if there was a model that could "dream" videos in real time, videos that are playable and personalized for the viewer, with adjustment knobs to make it fun, serene, or epic? We are one to two weeks away from our next-generation product launch.
04 Opinionated Viewpoints
1. Tencent Research Institute Roundtable Transcript: Amidst AI development, human imagination becomes the final unique advantage
Recently, the Tencent Research Institute released the latest episode of "Midsummer Six Talks," themed "How to Turn Imagination into a Competitive Advantage in the AI Era?" The guests in this dialogue included several founders of AI startups, who discussed the mutual transformation of imagination and competition in the AI era.
The host, Tencent Research Institute senior expert Yuan Xiaohui, pointed out that in an era where AI is increasingly capable of action, human imagination might become the final unique advantage. Looking ahead to the next 3 to 5 years:
- Neta CEO Hu Xiuyu predicts that individual creators will use agent IPs to build "one-person companies";
- Jingying Technology founder Zhu Jiang emphasizes that AI allows everyone to express their inner stories and interact more deeply with content;
- Tezign CEO Fan Ling believes agents will evolve from auxiliary tools into "digital employees" that can independently deliver results, and tool companies will also transform into result-oriented agent platforms;
- Reach Future co-founder You Wei focuses on the co-evolution of user habits and AI product forms, pointing out that the real opportunity lies not in the model capability itself, but in how to integrate it into specific usage scenarios and form new user collaboration habits.
Regarding the deeper issue of "whether AI will replace human subjectivity," the guests held sober judgments and complex emotions:
- Zhu Jiang believes the ultimate form of AI entertainment might be a more immersive content experience, where creation and consumption will tend to merge in the future.
- Hu Xiuyu admitted his concern that overly powerful technology could weaken equality in the cultural sphere, especially when AI not only undertakes creation but also judgment, potentially diluting human creative desire and the right to critique.
- Fan Ling finds confidence in history: just as the birth of the camera did not end art but instead opened up new art forms like abstraction, installation, and photography, AI may not be the end of human imagination, but another starting point. (@APPSO)
More Voice Agent Study Notes:
Written at the end:
We welcome more partners to participate in the co-creation of the 「RTE Developer Daily」 . Interested friends, please contact us through the developer community or public account messages, and remember the code word "co-creation".
For any feedback (including but not limited to content and format), we are extremely grateful and have small surprises in return. For example, what content you hope to see in the daily; your own recommended sources, projects, topics, events, etc.; or list a few content channels you like and often read; areas where content layout or presentation can be improved, etc.
Source materials from official media/online news