Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs

Adding a second GPU to a language-model server divides model work and memory but also introduces synchronization, collective communication, scheduling, and process overhead. We characterize the conditions under which two-way tensor parallelism (TP2) improves upstream vLLM serving on a pinned, PHB-connected pair of NVIDIA Tesla T4 GPUs in Kaggle. The principal experiment compares TP1 and TP2 across four serving-ready model families, three exact-token workload shapes, six concurrency levels, and five matched logical repetitions per model-workload combination. Of 60 terminal logical shards and 720 planned fresh-server cells, 55 shards yielded canonical performance outcomes and five terminated at a frozen resource boundary. The first sustained favorable output-throughput point under the predeclared paired-repetition 95% confidence-interval rule varied by model and workload. Prefill-heavy work crossed at concurrency 4 for Llama, Phi, and Ministral; their balanced and short workloads crossed later. Qwen crossed at concurrency 64 for balanced work, showed no sustained short-workload crossover, and resource-gated TP2 prefill-heavy work at concurrency 64. An independent two-rank NCCL microbenchmark characterizes communication on the same topology but does not isolate its causal contribution to serving. The kaggle-vllm toolkit delivers and records a checksum-pinned upstream runtime; it does not replace the inference engine. A separate ALLaM-7B Arabic/English case study illustrates descriptive high-concurrency throughput and context-capacity differences while retaining its weaker within-server uncertainty boundary. These findings characterize the tested runtime, workload, topology, and resource envelope rather than asserting a universal TP2 speedup. Associated software and artifacts: kaggle-vllm is an open-source compatibility and runtime-delivery toolkit for upstream vLLM on Kaggle dual NVIDIA Tesla T4 GPUs. Project source, releases, PyPI package, native CUDA runtime artifacts, and model artifacts are maintained separately.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23119478
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs

Waqas Mohammad
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs

Waqas Mohammad
preprint en

Abstract

Adding a second GPU to a language-model server divides model work and memory but also introduces synchronization, collective communication, scheduling, and process overhead. We characterize the conditions under which two-way tensor parallelism (TP2) improves upstream vLLM serving on a pinned, PHB-connected pair of NVIDIA Tesla T4 GPUs in Kaggle. The principal experiment compares TP1 and TP2 across four serving-ready model families, three exact-token workload shapes, six concurrency levels, and five matched logical repetitions per model-workload combination. Of 60 terminal logical shards and 720 planned fresh-server cells, 55 shards yielded canonical performance outcomes and five terminated at a frozen resource boundary. The first sustained favorable output-throughput point under the predeclared paired-repetition 95% confidence-interval rule varied by model and workload. Prefill-heavy work crossed at concurrency 4 for Llama, Phi, and Ministral; their balanced and short workloads crossed later. Qwen crossed at concurrency 64 for balanced work, showed no sustained short-workload crossover, and resource-gated TP2 prefill-heavy work at concurrency 64. An independent two-rank NCCL microbenchmark characterizes communication on the same topology but does not isolate its causal contribution to serving. The kaggle-vllm toolkit delivers and records a checksum-pinned upstream runtime; it does not replace the inference engine. A separate ALLaM-7B Arabic/English case study illustrates descriptive high-concurrency throughput and context-capacity differences while retaining its weaker within-server uncertainty boundary. These findings characterize the tested runtime, workload, topology, and resource envelope rather than asserting a universal TP2 speedup. Associated software and artifacts: kaggle-vllm is an open-source compatibility and runtime-delivery toolkit for upstream vLLM on Kaggle dual NVIDIA Tesla T4 GPUs. Project source, releases, PyPI package, native CUDA runtime artifacts, and model artifacts are maintained separately.

Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs — Waqas Mohammad · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS