By datathings
Run high-throughput LLM inference with vLLM: batch generation, chat completions, structured outputs, LoRA adapters, multimodal inputs, and embeddings via an OpenAI-compatible API.
Own this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimOwn this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimBased on adoption, maintenance, documentation, and repository signals. Not a security audit or endorsement.
Complete CBLAS and LAPACKE C API reference (LAPACK v3.12.1) covering 1284 functions across BLAS Level 1/2/3 vector and matrix operations, linear system solvers, eigenvalue/eigenvector computation, SVD, least squares, matrix factorizations, and auxiliary routines for scientific computing and high-performance numerical linear algebra.
OpenCL SDK (Khronos Group) v2026.05.29 (OpenCL 3.1) C/C++ skill — cross-platform GPU/CPU parallel computing with ~60 C API functions, C++ wrapper, and SDK utilities
ggml v0.15.3 C tensor library skill — 650+ functions and 98 ops for graph computation, GGUF I/O, multi-backend inference, GLU activations, SSM/RWKV ops, and ML training
pandapower v3.4.0 Python skill - power systems analysis with 80+ functions for AC/DC power flow, OPF, short circuit (IEC 60909), and state estimation
AMD ROCm 7.2.4 GPU computing stack: HIP kernel development, rocBLAS/rocFFT/rocRAND/rocSOLVER compute libraries, profiling, and CUDA-to-HIP porting
npx claudepluginhub datathings/marketplace --plugin vllmRun and manage local LLMs via Ollama's REST API, with support for text generation, chat, embeddings, and custom model creation.
Agent Skills for Together AI platform — inference, training, embeddings, audio, video, images, function calling, and infrastructure
Deploy and benchmark vLLM with Claude Code
Inference-time scaling for LLMs — generate multiple candidates and select the best using voting, scoring, or search
Agent-ready playbooks for LLM serving benchmarks, capacity planning, torch-profiler triage, pipeline analysis, compute simulation, SGLang/vLLM SOTA Humanize loops, human code review, production incident triage, and model PR-history dossiers.
When setting up local LLM inference without cloud APIs. When running GGUF models locally. When needing OpenAI-compatible API from a local model. When building offline/air-gapped AI tools. When troubleshooting local LLM server connections.