测试环境
机器: NVIDIA A100-SXM4-80GB
模型: Qwen3-30B-A3B
TP + EP
启动命令
CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008 --tp-size 2 --ep-size 2
压测命令
python3 -m sglang.bench_serving --backend sglang --dataset-name random --num-prompts 100 --max-concurrency 2 --random-input 512 --random-output 128 --random-range-ratio 0.75 --host 127.0.0.1 --port 3008 --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 48.20
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11313
Request throughput (req/s): 2.07
Input token throughput (tok/s): 920.38
Output token throughput (tok/s): 234.69
Peak output token throughput (tok/s): 246.00
Peak concurrent requests: 5
Total token throughput (tok/s): 1155.08
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 958.39
Median E2E Latency (ms): 960.76
P90 E2E Latency (ms): 1051.75
P99 E2E Latency (ms): 1379.85
---------------Time to First Token----------------
Mean TTFT (ms): 73.13
Median TTFT (ms): 70.07
P99 TTFT (ms): 185.89
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 7.89
Median TPOT (ms): 7.88
P99 TPOT (ms): 10.52
---------------Inter-Token Latency----------------
Mean ITL (ms): 7.89
Median ITL (ms): 7.39
P95 ITL (ms): 7.63
P99 ITL (ms): 10.58
Max ITL (ms): 421.28
==================================================
--max-concurrency 16
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 100
Benchmark duration (s): 99.93
Total input tokens: 180572
Total input text tokens: 180572
Total generated tokens: 90710
Total generated tokens (retokenized): 90710
Request throughput (req/s): 1.00
Input token throughput (tok/s): 1807.05
Output token throughput (tok/s): 907.77
Peak output token throughput (tok/s): 1200.00
Peak concurrent requests: 21
Total token throughput (tok/s): 2714.82
Concurrency: 15.04
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 15030.39
Median E2E Latency (ms): 15224.44
P90 E2E Latency (ms): 16900.46
P99 E2E Latency (ms): 17575.99
---------------Time to First Token----------------
Mean TTFT (ms): 198.91
Median TTFT (ms): 102.04
P99 TTFT (ms): 978.44
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 16.37
Median TPOT (ms): 16.82
P99 TPOT (ms): 17.32
---------------Inter-Token Latency----------------
Mean ITL (ms): 16.37
Median ITL (ms): 15.76
P95 ITL (ms): 16.60
P99 ITL (ms): 67.54
Max ITL (ms): 795.54
==================================================
TP + DP
启动命令
CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008 --tp-size 2 --dp-size 2
压测命令
python3 -m sglang.bench_serving --backend sglang --dataset-name random --num-prompts 100 --max-concurrency 2 --random-input 512 --random-output 128 --random-range-ratio 0.75 --host 127.0.0.1 --port 3008 --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 35.73
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11313
Request throughput (req/s): 2.80
Input token throughput (tok/s): 1241.57
Output token throughput (tok/s): 316.59
Peak output token throughput (tok/s): 334.00
Peak concurrent requests: 6
Total token throughput (tok/s): 1558.16
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 709.65
Median E2E Latency (ms): 704.56
P90 E2E Latency (ms): 784.77
P99 E2E Latency (ms): 980.85
---------------Time to First Token----------------
Mean TTFT (ms): 62.35
Median TTFT (ms): 57.40
P99 TTFT (ms): 130.09
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 5.77
Median TPOT (ms): 5.70
P99 TPOT (ms): 6.91
---------------Inter-Token Latency----------------
Mean ITL (ms): 5.77
Median ITL (ms): 5.70
P95 ITL (ms): 5.85
P99 ITL (ms): 6.81
Max ITL (ms): 254.38
==================================================
--max-concurrency 16
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 100
Benchmark duration (s): 71.09
Total input tokens: 180572
Total input text tokens: 180572
Total generated tokens: 90710
Total generated tokens (retokenized): 90710
Request throughput (req/s): 1.41
Input token throughput (tok/s): 2539.89
Output token throughput (tok/s): 1275.91
Peak output token throughput (tok/s): 1584.00
Peak concurrent requests: 23
Total token throughput (tok/s): 3815.80
Concurrency: 15.01
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 10667.88
Median E2E Latency (ms): 10844.28
P90 E2E Latency (ms): 11812.37
P99 E2E Latency (ms): 12531.83
---------------Time to First Token----------------
Mean TTFT (ms): 140.89
Median TTFT (ms): 90.60
P99 TTFT (ms): 522.70
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 11.63
Median TPOT (ms): 11.75
P99 TPOT (ms): 12.60
---------------Inter-Token Latency----------------
Mean ITL (ms): 11.62
Median ITL (ms): 11.23
P95 ITL (ms): 13.16
P99 ITL (ms): 27.40
Max ITL (ms): 366.62
==================================================
纯TP
启动命令
CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008 --tp-size 4
压测命令
python3 -m sglang.bench_serving --backend sglang --dataset-name random --num-prompts 100 --max-concurrency 2 --random-input 512 --random-output 128 --random-range-ratio 0.75 --host 127.0.0.1 --port 3008 --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 36.63
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11313
Request throughput (req/s): 2.73
Input token throughput (tok/s): 1211.19
Output token throughput (tok/s): 308.84
Peak output token throughput (tok/s): 331.00
Peak concurrent requests: 6
Total token throughput (tok/s): 1520.03
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 728.40
Median E2E Latency (ms): 735.26
P90 E2E Latency (ms): 807.11
P99 E2E Latency (ms): 859.59
---------------Time to First Token----------------
Mean TTFT (ms): 67.03
Median TTFT (ms): 64.15
P99 TTFT (ms): 173.31
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 5.90
Median TPOT (ms): 5.92
P99 TPOT (ms): 6.38
---------------Inter-Token Latency----------------
Mean ITL (ms): 5.90
Median ITL (ms): 5.45
P95 ITL (ms): 5.66
P99 ITL (ms): 8.40
Max ITL (ms): 70.97
==================================================
单TP
启动命令
CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008
压测命令
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 54.86
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11312
Request throughput (req/s): 1.82
Input token throughput (tok/s): 808.69
Output token throughput (tok/s): 206.21
Peak output token throughput (tok/s): 218.00
Peak concurrent requests: 5
Total token throughput (tok/s): 1014.90
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 1091.20
Median E2E Latency (ms): 1097.94
P90 E2E Latency (ms): 1217.14
P99 E2E Latency (ms): 1255.47
---------------Time to First Token----------------
Mean TTFT (ms): 70.57
Median TTFT (ms): 68.54
P99 TTFT (ms): 111.12
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 9.10
Median TPOT (ms): 9.18
P99 TPOT (ms): 9.62
---------------Inter-Token Latency----------------
Mean ITL (ms): 9.10
Median ITL (ms): 8.73
P95 ITL (ms): 9.02
P99 ITL (ms): 24.49
Max ITL (ms): 62.99
==================================================
纯DP
启动命令
CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008 --dp-size 4
压测命令
python3 -m sglang.bench_serving --backend sglang --dataset-name random --num-prompts 100 --max-concurrency 2 --random-input 512 --random-output 128 --random-range-ratio 0.75 --host 127.0.0.1 --port 3008 --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 40.71
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11313
Request throughput (req/s): 2.46
Input token throughput (tok/s): 1089.91
Output token throughput (tok/s): 277.92
Peak output token throughput (tok/s): 285.00
Peak concurrent requests: 6
Total token throughput (tok/s): 1367.83
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 808.03
Median E2E Latency (ms): 810.06
P90 E2E Latency (ms): 886.16
P99 E2E Latency (ms): 908.02
---------------Time to First Token----------------
Mean TTFT (ms): 60.12
Median TTFT (ms): 60.06
P99 TTFT (ms): 68.02
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 6.67
Median TPOT (ms): 6.67
P99 TPOT (ms): 6.77
---------------Inter-Token Latency----------------
Mean ITL (ms): 6.67
Median ITL (ms): 6.66
P95 ITL (ms): 6.80
P99 ITL (ms): 7.28
Max ITL (ms): 9.72
==================================================
--max-concurrency 16
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 100
Benchmark duration (s): 78.34
Total input tokens: 180572
Total input text tokens: 180572
Total generated tokens: 90710
Total generated tokens (retokenized): 90710
Request throughput (req/s): 1.28
Input token throughput (tok/s): 2304.83
Output token throughput (tok/s): 1157.83
Peak output token throughput (tok/s): 1404.00
Peak concurrent requests: 22
Total token throughput (tok/s): 3462.66
Concurrency: 14.94
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 11707.81
Median E2E Latency (ms): 11848.09
P90 E2E Latency (ms): 13279.15
P99 E2E Latency (ms): 14259.69
---------------Time to First Token----------------
Mean TTFT (ms): 163.18
Median TTFT (ms): 131.64
P99 TTFT (ms): 442.14
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 12.75
Median TPOT (ms): 12.79
P99 TPOT (ms): 14.77
---------------Inter-Token Latency----------------
Mean ITL (ms): 12.74
Median ITL (ms): 12.33
P95 ITL (ms): 15.95
P99 ITL (ms): 17.12
Max ITL (ms): 309.97
==================================================
加Nsight
启动命令
nsys profile --trace=cuda,nvtx --cudabacktrace=false --force-overwrite=true --output sglang_server_profile python -m sglang.launch_server --model-path /nfs/models/Qwen/Qwen3-30B-A3B --port 3008 --tp-size 2 --ep-size 2
压测命令
python3 -m sglang.bench_serving --backend sglang --dataset-name random --num-prompts 100 --max-concurrency 2 --random-input 512 --random-output 128 --random-range-ratio 0.75 --host 127.0.0.1 --port 3008 --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
数据
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 2
Successful requests: 100
Benchmark duration (s): 53.23
Total input tokens: 44366
Total input text tokens: 44366
Total generated tokens: 11313
Total generated tokens (retokenized): 11313
Request throughput (req/s): 1.88
Input token throughput (tok/s): 833.48
Output token throughput (tok/s): 212.53
Peak output token throughput (tok/s): 231.00
Peak concurrent requests: 4
Total token throughput (tok/s): 1046.01
Concurrency: 1.99
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 1058.26
Median E2E Latency (ms): 1073.28
P90 E2E Latency (ms): 1165.43
P99 E2E Latency (ms): 1231.51
---------------Time to First Token----------------
Mean TTFT (ms): 83.73
Median TTFT (ms): 82.14
P99 TTFT (ms): 157.14
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 8.69
Median TPOT (ms): 8.73
P99 TPOT (ms): 9.33
---------------Inter-Token Latency----------------
Mean ITL (ms): 8.69
Median ITL (ms): 8.15
P95 ITL (ms): 8.42
P99 ITL (ms): 13.75
Max ITL (ms): 87.77
==================================================