🎙️ X-AuT — Compressed Speech-LLM ASR
X-AuT progressively compresses the audio
encoder of Qwen3-ASR-0.6B from 18 to 14 transformer blocks (−20.7% audio-tower
parameters) and recovers the accuracy with cross-scale distillation from a frozen
Qwen3-ASR-1.7B teacher. Audio-in → transcript-out.
Upload or record audio (any common format, resampled to 16 kHz mono) and press
Transcribe.