Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs
Adding a second GPU to a language-model server divides model work and memory but also introduces synchronization, collective communication, scheduling, and process overhead. We characterize the conditions under which two-way tensor parallelism (TP2) improves upstream vLLM serving on a pinned, PHB-connected pair of NVIDIA Tesla T4 GPUs in Kaggle. The principal experiment compares TP1 and TP2 across four serving-ready model families, three exact-token workload shapes, six concurrency levels, and five matched logical repetitions per model-workload combination. Of 60 terminal logical shards and 720 planned fresh-server cells, 55 shards yielded canonical performance outcomes and five terminated at a frozen resource boundary. The first sustained favorable output-throughput point under the predeclared paired-repetition 95% confidence-interval rule varied by model and workload. Prefill-heavy work crossed at concurrency 4 for Llama, Phi, and Ministral; their balanced and short workloads crossed later. Qwen crossed at concurrency 64 for balanced work, showed no sustained short-workload crossover, and resource-gated TP2 prefill-heavy work at concurrency 64. An independent two-rank NCCL microbenchmark characterizes communication on the same topology but does not isolate its causal contribution to serving. The kaggle-vllm toolkit delivers and records a checksum-pinned upstream runtime; it does not replace the inference engine. A separate ALLaM-7B Arabic/English case study illustrates descriptive high-concurrency throughput and context-capacity differences while retaining its weaker within-server uncertainty boundary. These findings characterize the tested runtime, workload, topology, and resource envelope rather than asserting a universal TP2 speedup. Associated software and artifacts: kaggle-vllm is an open-source compatibility and runtime-delivery toolkit for upstream vLLM on Kaggle dual NVIDIA Tesla T4 GPUs. Project source, releases, PyPI package, native CUDA runtime artifacts, and model artifacts are maintained separately.
Authors
- Waqas Mohammad
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23119478
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint