Qwen3不同部署模型在A100设备实测数据

26 阅读7分钟

测试环境

机器: NVIDIA A100-SXM4-80GB

模型: Qwen3-30B-A3B

TP + EP

启动命令

CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server     --model-path /nfs/models/Qwen/Qwen3-30B-A3B    --port 3008 --tp-size 2 --ep-size 2 

压测命令

python3 -m sglang.bench_serving         --backend sglang         --dataset-name random         --num-prompts 100         --max-concurrency 2         --random-input 512         --random-output 128         --random-range-ratio 0.75         --host 127.0.0.1         --port 3008         --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json

数据

============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  48.20     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11313     
Request throughput (req/s):              2.07      
Input token throughput (tok/s):          920.38    
Output token throughput (tok/s):         234.69    
Peak output token throughput (tok/s):    246.00    
Peak concurrent requests:                5         
Total token throughput (tok/s):          1155.08   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   958.39    
Median E2E Latency (ms):                 960.76    
P90 E2E Latency (ms):                    1051.75   
P99 E2E Latency (ms):                    1379.85   
---------------Time to First Token----------------
Mean TTFT (ms):                          73.13     
Median TTFT (ms):                        70.07     
P99 TTFT (ms):                           185.89    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          7.89      
Median TPOT (ms):                        7.88      
P99 TPOT (ms):                           10.52     
---------------Inter-Token Latency----------------
Mean ITL (ms):                           7.89      
Median ITL (ms):                         7.39      
P95 ITL (ms):                            7.63      
P99 ITL (ms):                            10.58     
Max ITL (ms):                            421.28    
==================================================

--max-concurrency 16


============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 16        
Successful requests:                     100       
Benchmark duration (s):                  99.93     
Total input tokens:                      180572    
Total input text tokens:                 180572    
Total generated tokens:                  90710     
Total generated tokens (retokenized):    90710     
Request throughput (req/s):              1.00      
Input token throughput (tok/s):          1807.05   
Output token throughput (tok/s):         907.77    
Peak output token throughput (tok/s):    1200.00   
Peak concurrent requests:                21        
Total token throughput (tok/s):          2714.82   
Concurrency:                             15.04     
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   15030.39  
Median E2E Latency (ms):                 15224.44  
P90 E2E Latency (ms):                    16900.46  
P99 E2E Latency (ms):                    17575.99  
---------------Time to First Token----------------
Mean TTFT (ms):                          198.91    
Median TTFT (ms):                        102.04    
P99 TTFT (ms):                           978.44    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          16.37     
Median TPOT (ms):                        16.82     
P99 TPOT (ms):                           17.32     
---------------Inter-Token Latency----------------
Mean ITL (ms):                           16.37     
Median ITL (ms):                         15.76     
P95 ITL (ms):                            16.60     
P99 ITL (ms):                            67.54     
Max ITL (ms):                            795.54    
==================================================

TP + DP

启动命令

CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server     --model-path /nfs/models/Qwen/Qwen3-30B-A3B    --port 3008 --tp-size 2 --dp-size 2 

压测命令

python3 -m sglang.bench_serving         --backend sglang         --dataset-name random         --num-prompts 100         --max-concurrency 2         --random-input 512         --random-output 128         --random-range-ratio 0.75         --host 127.0.0.1         --port 3008         --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json

数据


============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  35.73     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11313     
Request throughput (req/s):              2.80      
Input token throughput (tok/s):          1241.57   
Output token throughput (tok/s):         316.59    
Peak output token throughput (tok/s):    334.00    
Peak concurrent requests:                6         
Total token throughput (tok/s):          1558.16   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   709.65    
Median E2E Latency (ms):                 704.56    
P90 E2E Latency (ms):                    784.77    
P99 E2E Latency (ms):                    980.85    
---------------Time to First Token----------------
Mean TTFT (ms):                          62.35     
Median TTFT (ms):                        57.40     
P99 TTFT (ms):                           130.09    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          5.77      
Median TPOT (ms):                        5.70      
P99 TPOT (ms):                           6.91      
---------------Inter-Token Latency----------------
Mean ITL (ms):                           5.77      
Median ITL (ms):                         5.70      
P95 ITL (ms):                            5.85      
P99 ITL (ms):                            6.81      
Max ITL (ms):                            254.38    
==================================================

--max-concurrency 16

============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 16        
Successful requests:                     100       
Benchmark duration (s):                  71.09     
Total input tokens:                      180572    
Total input text tokens:                 180572    
Total generated tokens:                  90710     
Total generated tokens (retokenized):    90710     
Request throughput (req/s):              1.41      
Input token throughput (tok/s):          2539.89   
Output token throughput (tok/s):         1275.91   
Peak output token throughput (tok/s):    1584.00   
Peak concurrent requests:                23        
Total token throughput (tok/s):          3815.80   
Concurrency:                             15.01     
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   10667.88  
Median E2E Latency (ms):                 10844.28  
P90 E2E Latency (ms):                    11812.37  
P99 E2E Latency (ms):                    12531.83  
---------------Time to First Token----------------
Mean TTFT (ms):                          140.89    
Median TTFT (ms):                        90.60     
P99 TTFT (ms):                           522.70    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          11.63     
Median TPOT (ms):                        11.75     
P99 TPOT (ms):                           12.60     
---------------Inter-Token Latency----------------
Mean ITL (ms):                           11.62     
Median ITL (ms):                         11.23     
P95 ITL (ms):                            13.16     
P99 ITL (ms):                            27.40     
Max ITL (ms):                            366.62    
==================================================

纯TP

启动命令

CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server     --model-path /nfs/models/Qwen/Qwen3-30B-A3B    --port 3008 --tp-size 4

压测命令

python3 -m sglang.bench_serving         --backend sglang         --dataset-name random         --num-prompts 100         --max-concurrency 2         --random-input 512         --random-output 128         --random-range-ratio 0.75         --host 127.0.0.1         --port 3008         --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json

数据

============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  36.63     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11313     
Request throughput (req/s):              2.73      
Input token throughput (tok/s):          1211.19   
Output token throughput (tok/s):         308.84    
Peak output token throughput (tok/s):    331.00    
Peak concurrent requests:                6         
Total token throughput (tok/s):          1520.03   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   728.40    
Median E2E Latency (ms):                 735.26    
P90 E2E Latency (ms):                    807.11    
P99 E2E Latency (ms):                    859.59    
---------------Time to First Token----------------
Mean TTFT (ms):                          67.03     
Median TTFT (ms):                        64.15     
P99 TTFT (ms):                           173.31    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          5.90      
Median TPOT (ms):                        5.92      
P99 TPOT (ms):                           6.38      
---------------Inter-Token Latency----------------
Mean ITL (ms):                           5.90      
Median ITL (ms):                         5.45      
P95 ITL (ms):                            5.66      
P99 ITL (ms):                            8.40      
Max ITL (ms):                            70.97     
==================================================

单TP

启动命令

CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server     --model-path /nfs/models/Qwen/Qwen3-30B-A3B    --port 3008

压测命令

数据


============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  54.86     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11312     
Request throughput (req/s):              1.82      
Input token throughput (tok/s):          808.69    
Output token throughput (tok/s):         206.21    
Peak output token throughput (tok/s):    218.00    
Peak concurrent requests:                5         
Total token throughput (tok/s):          1014.90   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   1091.20   
Median E2E Latency (ms):                 1097.94   
P90 E2E Latency (ms):                    1217.14   
P99 E2E Latency (ms):                    1255.47   
---------------Time to First Token----------------
Mean TTFT (ms):                          70.57     
Median TTFT (ms):                        68.54     
P99 TTFT (ms):                           111.12    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          9.10      
Median TPOT (ms):                        9.18      
P99 TPOT (ms):                           9.62      
---------------Inter-Token Latency----------------
Mean ITL (ms):                           9.10      
Median ITL (ms):                         8.73      
P95 ITL (ms):                            9.02      
P99 ITL (ms):                            24.49     
Max ITL (ms):                            62.99     
==================================================

纯DP

启动命令

CUDA_VISIBLE_DEVICES=4,5,6,7 python -m sglang.launch_server     --model-path /nfs/models/Qwen/Qwen3-30B-A3B    --port 3008 --dp-size 4

压测命令

python3 -m sglang.bench_serving         --backend sglang         --dataset-name random         --num-prompts 100         --max-concurrency 2         --random-input 512         --random-output 128         --random-range-ratio 0.75         --host 127.0.0.1         --port 3008         --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json

数据

============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  40.71     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11313     
Request throughput (req/s):              2.46      
Input token throughput (tok/s):          1089.91   
Output token throughput (tok/s):         277.92    
Peak output token throughput (tok/s):    285.00    
Peak concurrent requests:                6         
Total token throughput (tok/s):          1367.83   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   808.03    
Median E2E Latency (ms):                 810.06    
P90 E2E Latency (ms):                    886.16    
P99 E2E Latency (ms):                    908.02    
---------------Time to First Token----------------
Mean TTFT (ms):                          60.12     
Median TTFT (ms):                        60.06     
P99 TTFT (ms):                           68.02     
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          6.67      
Median TPOT (ms):                        6.67      
P99 TPOT (ms):                           6.77      
---------------Inter-Token Latency----------------
Mean ITL (ms):                           6.67      
Median ITL (ms):                         6.66      
P95 ITL (ms):                            6.80      
P99 ITL (ms):                            7.28      
Max ITL (ms):                            9.72      
==================================================

--max-concurrency 16

Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 16        
Successful requests:                     100       
Benchmark duration (s):                  78.34     
Total input tokens:                      180572    
Total input text tokens:                 180572    
Total generated tokens:                  90710     
Total generated tokens (retokenized):    90710     
Request throughput (req/s):              1.28      
Input token throughput (tok/s):          2304.83   
Output token throughput (tok/s):         1157.83   
Peak output token throughput (tok/s):    1404.00   
Peak concurrent requests:                22        
Total token throughput (tok/s):          3462.66   
Concurrency:                             14.94     
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   11707.81  
Median E2E Latency (ms):                 11848.09  
P90 E2E Latency (ms):                    13279.15  
P99 E2E Latency (ms):                    14259.69  
---------------Time to First Token----------------
Mean TTFT (ms):                          163.18    
Median TTFT (ms):                        131.64    
P99 TTFT (ms):                           442.14    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          12.75     
Median TPOT (ms):                        12.79     
P99 TPOT (ms):                           14.77     
---------------Inter-Token Latency----------------
Mean ITL (ms):                           12.74     
Median ITL (ms):                         12.33     
P95 ITL (ms):                            15.95     
P99 ITL (ms):                            17.12     
Max ITL (ms):                            309.97    
==================================================

加Nsight

启动命令

nsys profile     --trace=cuda,nvtx     --cudabacktrace=false     --force-overwrite=true     --output sglang_server_profile     python -m sglang.launch_server         --model-path /nfs/models/Qwen/Qwen3-30B-A3B         --port 3008         --tp-size 2         --ep-size 2

压测命令

python3 -m sglang.bench_serving         --backend sglang         --dataset-name random         --num-prompts 100         --max-concurrency 2         --random-input 512         --random-output 128         --random-range-ratio 0.75         --host 127.0.0.1         --port 3008         --dataset-path /nfs/datasets/ShareGPT_V3_unfiltered_cleaned_split.json

数据

============ Serving Benchmark Result ============
Backend:                                 sglang    
Traffic request rate:                    inf       
Max request concurrency:                 2         
Successful requests:                     100       
Benchmark duration (s):                  53.23     
Total input tokens:                      44366     
Total input text tokens:                 44366     
Total generated tokens:                  11313     
Total generated tokens (retokenized):    11313     
Request throughput (req/s):              1.88      
Input token throughput (tok/s):          833.48    
Output token throughput (tok/s):         212.53    
Peak output token throughput (tok/s):    231.00    
Peak concurrent requests:                4         
Total token throughput (tok/s):          1046.01   
Concurrency:                             1.99      
----------------End-to-End Latency----------------
Mean E2E Latency (ms):                   1058.26   
Median E2E Latency (ms):                 1073.28   
P90 E2E Latency (ms):                    1165.43   
P99 E2E Latency (ms):                    1231.51   
---------------Time to First Token----------------
Mean TTFT (ms):                          83.73     
Median TTFT (ms):                        82.14     
P99 TTFT (ms):                           157.14    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          8.69      
Median TPOT (ms):                        8.73      
P99 TPOT (ms):                           9.33      
---------------Inter-Token Latency----------------
Mean ITL (ms):                           8.69      
Median ITL (ms):                         8.15      
P95 ITL (ms):                            8.42      
P99 ITL (ms):                            13.75     
Max ITL (ms):                            87.77     
==================================================