Skills / Engineering / Groq Fast Inference

Groq Fast Inference

Run ultra-low-latency LLM inference on GroqCloud. Automate chat completions, audio transcription and translation, and text-to-speech voice management through Groq's high-performance API.

This skill integrates GroqCloud for fast inference. It calls chat completions on Groq-hosted models, transcribes and translates audio, manages TTS voices, and wires Groq into apps and pipelines that need very low latency and high throughput for real-time assistants and streaming responses.

groq inference llm low-latency speech

When to use

Use when you need very fast LLM inference, streaming chat, audio transcription/translation, or text-to-speech and want to run it on GroqCloud.

Examples

Low-latency chat endpoint

Stream completions from Groq

Build a streaming chat endpoint backed by GroqCloud so responses start rendering in under 300ms

Transcribe an audio file

Speech-to-text via Groq

Use Groq's audio API to transcribe a meeting recording and return timestamped segments
Added to wishlist