Groq Fast Inference
Run ultra-low-latency LLM inference on GroqCloud. Automate chat completions, audio transcription and translation, and text-to-speech voice management through Groq's high-performance API.
This skill integrates GroqCloud for fast inference. It calls chat completions on Groq-hosted models, transcribes and translates audio, manages TTS voices, and wires Groq into apps and pipelines that need very low latency and high throughput for real-time assistants and streaming responses.
When to use
Use when you need very fast LLM inference, streaming chat, audio transcription/translation, or text-to-speech and want to run it on GroqCloud.
Examples
Low-latency chat endpoint
Stream completions from Groq
Build a streaming chat endpoint backed by GroqCloud so responses start rendering in under 300ms
Transcribe an audio file
Speech-to-text via Groq
Use Groq's audio API to transcribe a meeting recording and return timestamped segments