(Realtime) Temporal Convolutions in PyTorch
-
Updated
Apr 7, 2025 - Python
(Realtime) Temporal Convolutions in PyTorch
Streamable Text-to-Speech model using a language modeling approach, without vector quantization
Dual-model speech AI toolkit for speaker verification and speaker-aware diarization, with streaming inference, meeting analysis, long-audio monitoring, and speaker-bank integration.
Pure PyTorch + 🤗 Transformers reimplementation of Megalodon (CEMA + chunked attention) - readable, hackable, no CUDA kernels required
Lossless AI model compression - ~34% smaller with bit-identical weights; the autopilot profiles your machine, picks the highest fidelity that runs, and streams models bigger than your RAM.
Open ML systems platform for training, profiling, evaluating, and serving AI models.
World's most deployable time series foundation model — 200K-6.5M params, zero-shot forecasting, streaming RNN inference, ONNX edge deployment, runs on Raspberry Pi
A defensive publication establishing dated public prior art over whole-answer buffering with delta-granular replay, a confidence-thresholded three-outcome dual-lane safety lattice, a provenance-complete abuse-harvest retraining lane, and an inverted
Efficient State Space Model layers in pure PyTorch — FFT training, streaming inference, ONNX export for edge deployment
CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs [FPL'26]
Real-time music-genre classification: spectrogram CNN, ONNX-optimised, served as a streaming/chunked classifier with PyTorch-vs-ONNX benchmarks. Track-aware GTZAN eval.
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
High-performance JAX-to-TensorRT compilation pipeline and decoupled gRPC streaming inference server for quantitative trading architectures.
文本 / 提示驱动的可控声音生成框架:情感 / 风格 / 音色可控,KV cache + 分类器无关引导(CFG),支持流式推理(纯 numpy 核心,可选 torch)
Real-time voice AI microservice - WebRTC, multi-tenant architecture, STT/TTS, streaming inference
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
Streaming version of S4ND-U-Net
Add a description, image, and links to the streaming-inference topic page so that developers can more easily learn about it.
To associate your repository with the streaming-inference topic, visit your repo's landing page and select "manage topics."