Zhipu AI released a new generation of speech recognition product series GLM-ASR, and simultaneously launched the desktop "Zhipu AI Input Method", which deeply integrates speech recognition with large model capabilities, focusing on the human-computer interaction method of "speaking and giving instructions". The GLM-ASR series includes the GLM-ASR-2512 cloud-side model and the GLM-ASR-Nano-2512 device-side model, covering a wide range of operating environments, from servers to notebooks and mobile phones.
According to reports, GLM-ASR-2512 is deployed in the cloud, with a character error rate of 0.0717, supports Chinese, English and some dialects, and is optimized for noisy environments, suitable for online service scenarios with high accuracy requirements. GLM-ASR-Nano-2512 has approximately 1.5 billion parameters and can run locally on end-side devices such as laptops and mobile phones, emphasizing low latency and real-time interactive experience, model weights and inference code are open-source, and the cloud version is provided through APIs.
"Zhipu AI Input Method" is aimed at desktop users, combining GLM-ASR with large models to support text input, application control, and command issuance through speech. The product introduces a "Whisper Mode" that recognizes speech even at lower volumes or breath, making it easy to use in quiet environments. Official reminder that when using voice input and continuous monitoring functions, users should pay attention to microphone permissions and data upload range, and configure them in combination with their own privacy and security requirements.
FAQs
Q: What is GLM-ASR?
A: GLM-ASR is a speech recognition model series launched by Zhipu AI, which is suitable for speech-to-text and voice control scenarios, including cloud and device-side versions.
Q: What is the difference between the GLM-ASR-2512 and the GLM-ASR-Nano-2512?
A: The former is deployed in the cloud, with higher accuracy, and is called through API; The latter has smaller parameters and can run locally on laptops, mobile phones, etc., focusing on low latency and offline availability.
Q: Is the GLM-ASR-Nano open source?
A: GLM-ASR-NANO-2512 provides open-source weighting and inference code, making it convenient for developers to integrate and develop in their own environment locally or in their own environment.
Q: What can "Zhipu AI Input Method" do?
A: This input method combines speech recognition with large models, supports voice input text, voice issuance of system instructions, and provides whisper mode to adapt to low-volume and privacy scenarios.
Q: What are the risks to be aware of when using these voice products?
A: Speech recognition involves microphone collection and possible data upload, and users should check app permissions, understand data usage rules, and carefully choose whether to enable cloud services or listen for a long time when sensitive information is involved.