Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339]
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339] █ █ █▄ ▄█
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.25.2.dev0+g752a3a504.d20260803
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339] █▄█▀ █ █ █ █ model poolside/Laguna-S-2.1-NVFP4
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:339]
- (APIServer pid=48) INFO 08-03 22:50:38 [api_utils.py:273] non-default args: {'model_tag': 'poolside/Laguna-S-2.1-NVFP4', 'default_chat_template_kwargs': {'enable_thinking': True}, 'enable_auto_tool_choice': True, 'tool_call_parser': 'poolside_v1', 'host': '0.0.0.0', 'port': 8888, 'model': 'poolside/Laguna-S-2.1-NVFP4', 'max_model_len': 262144, 'reasoning_parser': 'poolside_v1', 'master_addr': '192.168.9.21', 'master_port': 25000, 'nnodes': 2, 'tensor_parallel_size': 2, 'gpu_memory_utilization': 0.85, 'max_num_batched_tokens': 4096, 'max_num_seqs': 8}
- (APIServer pid=48) INFO 08-03 22:50:38 [arg_utils.py:772] HF_HUB_OFFLINE is True, replace model_id [poolside/Laguna-S-2.1-NVFP4] to model_path [/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669]
- (APIServer pid=48) WARNING 08-03 22:50:38 [envs.py:2041] Unknown vLLM environment variable detected: VLLM_BASE_DIR
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) INFO 08-03 22:50:43 [model.py:619] Resolved architecture: LagunaForCausalLM
- (APIServer pid=48) INFO 08-03 22:50:43 [model.py:1776] Using max model len 262144
- (APIServer pid=48) INFO 08-03 22:50:44 [arg_utils.py:2026] Inferred data_parallel_rank 0 from node_rank 0
- (APIServer pid=48) INFO 08-03 22:50:44 [scheduler.py:252] Chunked prefill is enabled with max_num_batched_tokens=4096.
- (APIServer pid=48) INFO 08-03 22:50:44 [vllm.py:1042] Asynchronous scheduling is enabled.
- (APIServer pid=48) INFO 08-03 22:50:44 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] The tokenizer you are loading from '/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
- (APIServer pid=48) WARNING 08-03 22:50:44 [vllm.py:1486] Auto-initialization of reasoning token IDs failed. Please check whether your reasoning parser has implemented the `reasoning_start_str` and `reasoning_end_str`.
- (APIServer pid=48) INFO 08-03 22:50:44 [compilation.py:312] Enabled custom fusions: act_quant
- (EngineCore pid=139) INFO 08-03 22:50:48 [core.py:114] Initializing a V1 LLM engine (v0.25.2.dev0+g752a3a504.d20260803) with config: model='/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669', speculative_config=None, tokenizer='/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=262144, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=compressed-tensors, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='poolside_v1', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.VLLM_COMPILE: 3>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::qwen_gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::hpc_rope_norm_forward', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [4096], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 16, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
- (EngineCore pid=139) INFO 08-03 22:50:48 [multiproc_executor.py:140] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=192.168.9.21, mq_connect_ip=192.168.9.21 (local), world_size=2, local_world_size=1
- (Worker pid=165) INFO 08-03 22:50:52 [parallel_state.py:1607] world_size=2 rank=0 local_rank=0 distributed_init_method=tcp://192.168.9.21:25000 backend=nccl
- (Worker pid=165) INFO 08-03 22:51:19 [pynccl.py:113] vLLM is using nccl==2.30.7
- (Worker pid=165) WARNING 08-03 22:51:21 [symm_mem.py:66] SymmMemCommunicator: Device capability 12.1 not supported, communicator is not available.
- (Worker pid=165) INFO 08-03 22:51:21 [cuda_communicator.py:264] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'tp:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
- (Worker pid=165) INFO 08-03 22:51:23 [cuda_communicator.py:264] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'ep:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
- (Worker pid=165) INFO 08-03 22:51:23 [parallel_state.py:1942] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
- (Worker pid=165) INFO 08-03 22:51:23 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
- (Worker_TP0 pid=165) INFO 08-03 22:51:23 [gpu_model_runner.py:5209] Starting to load model /cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669...
- (Worker_TP0 pid=165) INFO 08-03 22:51:24 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
- (Worker_TP0 pid=165) INFO 08-03 22:51:24 [flash_attn.py:718] Using FlashAttention version 2
- (Worker_TP0 pid=165) INFO 08-03 22:51:24 [nvfp4.py:285] Using 'FLASHINFER_CUTLASS' NvFp4 MoE backend out of potential backends: ['FLASHINFER_TRTLLM', 'FLASHINFER_CUTEDSL', 'FLASHINFER_CUTEDSL_BATCHED', 'FLASHINFER_CUTLASS', 'VLLM_CUTLASS', 'MARLIN', 'HUMMING', 'EMULATION'].
- (Worker_TP0 pid=165) INFO 08-03 22:51:25 [unquantized.py:262] Using FlashInfer CUTLASS Unquantized MoE backend out of potential backends: ['FlashInfer TRTLLM', 'FlashInfer CUTLASS', 'TRITON', 'BATCHED_TRITON'].
- (Worker_TP0 pid=165) INFO 08-03 22:51:26 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 92.85 GiB. Available RAM: 59.51 GiB.
- (Worker_TP0 pid=165) INFO 08-03 22:51:26 [weight_utils.py:879] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre) and the checkpoint size (92.85 GiB) exceeds 90% of available RAM (59.51 GiB).
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 0% Completed | 0/49 [00:00<?, ?it/s]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 2% Completed | 1/49 [00:00<00:36, 1.33it/s]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 4% Completed | 2/49 [00:02<00:52, 1.12s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 6% Completed | 3/49 [00:04<01:16, 1.66s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 8% Completed | 4/49 [00:19<05:21, 7.14s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 10% Completed | 5/49 [00:36<07:50, 10.69s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 12% Completed | 6/49 [00:54<09:17, 12.96s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 14% Completed | 7/49 [01:11<10:06, 14.44s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 16% Completed | 8/49 [01:27<10:08, 14.83s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 18% Completed | 9/49 [01:44<10:15, 15.39s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 20% Completed | 10/49 [02:00<10:17, 15.84s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 22% Completed | 11/49 [02:12<09:08, 14.44s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 24% Completed | 12/49 [02:20<07:46, 12.61s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 27% Completed | 13/49 [02:28<06:37, 11.06s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 29% Completed | 14/49 [02:36<05:54, 10.14s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 31% Completed | 15/49 [02:44<05:22, 9.49s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 33% Completed | 16/49 [02:51<04:53, 8.90s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 35% Completed | 17/49 [02:59<04:37, 8.69s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 37% Completed | 18/49 [03:07<04:18, 8.32s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 39% Completed | 19/49 [03:15<04:08, 8.28s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 41% Completed | 20/49 [03:23<04:00, 8.28s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 43% Completed | 21/49 [03:31<03:51, 8.26s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 45% Completed | 22/49 [03:40<03:42, 8.23s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 47% Completed | 23/49 [03:47<03:29, 8.06s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 49% Completed | 24/49 [03:56<03:22, 8.12s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 51% Completed | 25/49 [04:04<03:14, 8.10s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 53% Completed | 26/49 [04:12<03:06, 8.12s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 55% Completed | 27/49 [04:20<02:59, 8.15s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 57% Completed | 28/49 [04:28<02:52, 8.21s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 59% Completed | 29/49 [04:36<02:42, 8.14s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 61% Completed | 30/49 [04:44<02:33, 8.10s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 63% Completed | 31/49 [04:53<02:26, 8.11s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 65% Completed | 32/49 [05:00<02:16, 8.04s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 67% Completed | 33/49 [05:08<02:07, 7.97s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 69% Completed | 34/49 [05:16<01:59, 7.95s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 71% Completed | 35/49 [05:24<01:51, 7.97s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 73% Completed | 36/49 [05:32<01:44, 8.02s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 76% Completed | 37/49 [05:40<01:36, 8.01s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 78% Completed | 38/49 [05:48<01:27, 7.93s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 80% Completed | 39/49 [05:56<01:19, 7.94s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 82% Completed | 40/49 [06:04<01:12, 8.01s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 84% Completed | 41/49 [06:12<01:03, 7.99s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 86% Completed | 42/49 [06:20<00:55, 7.99s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 88% Completed | 43/49 [06:27<00:47, 7.84s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 90% Completed | 44/49 [06:35<00:39, 7.83s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 92% Completed | 45/49 [06:43<00:31, 7.88s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 94% Completed | 46/49 [06:51<00:23, 7.69s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 96% Completed | 47/49 [06:58<00:15, 7.55s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 98% Completed | 48/49 [07:05<00:07, 7.44s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 100% Completed | 49/49 [07:08<00:00, 6.17s/it]
- (Worker_TP0 pid=165)
- Loading safetensors checkpoint shards: 100% Completed | 49/49 [07:08<00:00, 8.75s/it]
- (Worker_TP0 pid=165)
- (Worker_TP0 pid=165) INFO 08-03 22:58:35 [default_loader.py:430] Loading weights took 428.72 seconds
- (Worker_TP0 pid=165) INFO 08-03 22:58:35 [nvfp4.py:543] Using MoEPrepareAndFinalizeNoDPEPModular
- (Worker_TP0 pid=165) INFO 08-03 22:58:35 [unquantized.py:334] Using MoEPrepareAndFinalizeNoDPEPModular
- (Worker_TP0 pid=165) INFO 08-03 22:58:36 [gpu_model_runner.py:5306] Model loading took 46.91 GiB memory and 431.829826 seconds
- (Worker_TP0 pid=165) INFO 08-03 22:58:42 [backends.py:1089] Using cache directory: /root/.cache/vllm/torch_compile_cache/8e143b869f/rank_0_0/backbone for vLLM's torch.compile
- (Worker_TP0 pid=165) INFO 08-03 22:58:42 [backends.py:1148] Dynamo bytecode transform time: 6.39 s
- (Worker_TP0 pid=165) [rank0]:W0803 22:58:44.237000 165 torch/_inductor/utils.py:1731] Not enough SMs to use max_autotune_gemm mode
- (Worker_TP0 pid=165) INFO 08-03 22:58:49 [backends.py:378] Cache the graph of compile range (1, 4096) for later use
- (Worker_TP0 pid=165) INFO 08-03 22:59:05 [backends.py:393] Compiling a graph for compile range (1, 4096) takes 21.96 s
- (Worker_TP0 pid=165) INFO 08-03 22:59:10 [decorators.py:708] saved AOT compiled function to /root/.cache/vllm/torch_compile_cache/torch_aot_compile/7d47e56b466554a017361f3961e5e87470c9ad4f5c053d61be339aca75603d67/rank_0_0/model
- (Worker_TP0 pid=165) INFO 08-03 22:59:10 [monitor.py:53] torch.compile took 34.42 s in total
- (Worker_TP0 pid=165) INFO 08-03 22:59:14 [monitor.py:81] Initial profiling/warmup run took 3.37 s
- (Worker_TP0 pid=165) INFO 08-03 22:59:19 [gpu_model_runner.py:6534] Profiling CUDA graph memory: PIECEWISE=5 (largest=16), FULL=4 (largest=8)
- (EngineCore pid=139) INFO 08-03 22:59:37 [shm_broadcast.py:705] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
- (Worker_TP0 pid=165) INFO 08-03 23:00:09 [gpu_model_runner.py:6639] Estimated CUDA graph memory: 0.90 GiB total
- (Worker_TP0 pid=165) INFO 08-03 23:00:10 [gpu_worker.py:531] Freed 0.00 GiB before KV cache sizing; non-torch profile increase is 3.82 GiB.
- (Worker_TP0 pid=165) INFO 08-03 23:00:10 [gpu_worker.py:569] Available KV cache memory: 51.1 GiB
- (Worker_TP0 pid=165) INFO 08-03 23:00:10 [gpu_worker.py:584] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.8500 is equivalent to --gpu-memory-utilization=0.8426 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.8574. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0.
- (EngineCore pid=139) INFO 08-03 23:00:10 [kv_cache_utils.py:2146] GPU KV cache size: 2,113,609 tokens
- (EngineCore pid=139) INFO 08-03 23:00:10 [kv_cache_utils.py:2147] Maximum concurrency for 262,144 tokens per request: 8.06x
- (Worker_TP0 pid=165) INFO 08-03 23:00:11 [gpu_worker.py:739] Cleared 0.15 GiB of cached CUDA allocator memory before KV cache allocation.
- (Worker_TP0 pid=165) INFO 08-03 23:00:26 [deep_gemm.py:175] deep_gemm not found in site-packages, trying vendored vllm.third_party.deep_gemm
- (Worker_TP0 pid=165) INFO 08-03 23:00:26 [deep_gemm.py:202] DeepGEMM PDL enabled on vllm.third_party.deep_gemm.
- (Worker_TP0 pid=165) 2026-08-03 23:00:26,621 - INFO - autotuner.py:829 - flashinfer.jit: [Autotuner]: Autotuning process starts ...
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 0%| | 0/21 [00:00<?, ?profile/s]2026-08-03 23:00:30,567 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 5%|▍ | 1/21 [00:03<01:09, 3.45s/profile]2026-08-03 23:00:30,648 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165) 2026-08-03 23:00:30,752 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 14%|█▍ | 3/21 [00:03<00:17, 1.04profile/s]2026-08-03 23:00:30,908 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 19%|█▉ | 4/21 [00:03<00:11, 1.46profile/s]2026-08-03 23:00:31,158 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 24%|██▍ | 5/21 [00:04<00:08, 1.85profile/s]2026-08-03 23:00:31,525 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 29%|██▊ | 6/21 [00:04<00:07, 2.06profile/s]2026-08-03 23:00:31,985 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 33%|███▎ | 7/21 [00:04<00:06, 2.10profile/s]2026-08-03 23:00:32,490 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 38%|███▊ | 8/21 [00:05<00:06, 2.06profile/s]2026-08-03 23:00:32,998 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 43%|████▎ | 9/21 [00:05<00:05, 2.03profile/s]2026-08-03 23:00:33,530 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 48%|████▊ | 10/21 [00:06<00:05, 1.98profile/s]2026-08-03 23:00:34,068 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 52%|█████▏ | 11/21 [00:06<00:05, 1.94profile/s]2026-08-03 23:00:34,614 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 57%|█████▋ | 12/21 [00:07<00:04, 1.91profile/s]2026-08-03 23:00:35,192 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 62%|██████▏ | 13/21 [00:08<00:04, 1.85profile/s]2026-08-03 23:00:35,782 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 67%|██████▋ | 14/21 [00:08<00:03, 1.80profile/s]2026-08-03 23:00:36,398 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 71%|███████▏ | 15/21 [00:09<00:03, 1.74profile/s]2026-08-03 23:00:37,010 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 76%|███████▌ | 16/21 [00:09<00:02, 1.71profile/s]2026-08-03 23:00:37,660 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 81%|████████ | 17/21 [00:10<00:02, 1.65profile/s]2026-08-03 23:00:38,336 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 86%|████████▌ | 18/21 [00:11<00:01, 1.60profile/s]2026-08-03 23:00:39,077 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 90%|█████████ | 19/21 [00:11<00:01, 1.51profile/s]2026-08-03 23:00:39,867 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 95%|█████████▌| 20/21 [00:12<00:00, 1.43profile/s]2026-08-03 23:00:41,116 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 4 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 100%|██████████| 21/21 [00:14<00:00, 1.16profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 100%|██████████| 21/21 [00:14<00:00, 1.50profile/s]
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 0%| | 0/21 [00:00<?, ?profile/s]2026-08-03 23:00:41,333 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 5%|▍ | 1/21 [00:00<00:04, 4.78profile/s]2026-08-03 23:00:41,489 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 10%|▉ | 2/21 [00:00<00:03, 5.61profile/s]2026-08-03 23:00:41,670 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 14%|█▍ | 3/21 [00:00<00:03, 5.58profile/s]2026-08-03 23:00:41,887 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 19%|█▉ | 4/21 [00:00<00:03, 5.15profile/s]2026-08-03 23:00:42,205 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 24%|██▍ | 5/21 [00:01<00:03, 4.19profile/s]2026-08-03 23:00:42,639 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 29%|██▊ | 6/21 [00:01<00:04, 3.27profile/s]2026-08-03 23:00:43,153 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 33%|███▎ | 7/21 [00:02<00:05, 2.68profile/s]2026-08-03 23:00:43,698 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 38%|███▊ | 8/21 [00:02<00:05, 2.34profile/s]2026-08-03 23:00:44,243 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 43%|████▎ | 9/21 [00:03<00:05, 2.15profile/s]2026-08-03 23:00:44,828 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 48%|████▊ | 10/21 [00:03<00:05, 1.99profile/s]2026-08-03 23:00:45,468 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 52%|█████▏ | 11/21 [00:04<00:05, 1.84profile/s]2026-08-03 23:00:46,147 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 57%|█████▋ | 12/21 [00:05<00:05, 1.71profile/s]2026-08-03 23:00:46,860 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 62%|██████▏ | 13/21 [00:05<00:04, 1.60profile/s]2026-08-03 23:00:47,629 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 67%|██████▋ | 14/21 [00:06<00:04, 1.50profile/s]2026-08-03 23:00:48,442 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 71%|███████▏ | 15/21 [00:07<00:04, 1.41profile/s]2026-08-03 23:00:49,341 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 76%|███████▌ | 16/21 [00:08<00:03, 1.30profile/s]2026-08-03 23:00:50,319 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 81%|████████ | 17/21 [00:09<00:03, 1.20profile/s]2026-08-03 23:00:51,426 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 86%|████████▌ | 18/21 [00:10<00:02, 1.09profile/s]2026-08-03 23:00:52,818 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 90%|█████████ | 19/21 [00:11<00:02, 1.06s/profile]2026-08-03 23:00:54,262 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 95%|█████████▌| 20/21 [00:13<00:01, 1.17s/profile]2026-08-03 23:00:57,085 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 10 unsupported tactic(s) for trtllm::fused_moe::gemm2 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 100%|██████████| 21/21 [00:15<00:00, 1.67s/profile]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 100%|██████████| 21/21 [00:15<00:00, 1.32profile/s]
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 0%| | 0/21 [00:00<?, ?profile/s]2026-08-03 23:00:58,880 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 5%|▍ | 1/21 [00:00<00:10, 1.94profile/s]2026-08-03 23:00:58,933 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165) 2026-08-03 23:00:59,024 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 14%|█▍ | 3/21 [00:00<00:03, 5.34profile/s]2026-08-03 23:00:59,185 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 19%|█▉ | 4/21 [00:00<00:03, 5.61profile/s]2026-08-03 23:00:59,453 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 24%|██▍ | 5/21 [00:01<00:03, 4.81profile/s]2026-08-03 23:00:59,869 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 29%|██▊ | 6/21 [00:01<00:04, 3.64profile/s]2026-08-03 23:01:00,392 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 33%|███▎ | 7/21 [00:02<00:04, 2.83profile/s]2026-08-03 23:01:00,971 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 38%|███▊ | 8/21 [00:02<00:05, 2.37profile/s]2026-08-03 23:01:01,556 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 43%|████▎ | 9/21 [00:03<00:05, 2.12profile/s]2026-08-03 23:01:02,171 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 48%|████▊ | 10/21 [00:03<00:05, 1.94profile/s]2026-08-03 23:01:02,799 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 52%|█████▏ | 11/21 [00:04<00:05, 1.82profile/s]2026-08-03 23:01:03,438 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 57%|█████▋ | 12/21 [00:05<00:05, 1.73profile/s]2026-08-03 23:01:04,096 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 62%|██████▏ | 13/21 [00:05<00:04, 1.66profile/s]2026-08-03 23:01:04,811 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 67%|██████▋ | 14/21 [00:06<00:04, 1.57profile/s]2026-08-03 23:01:05,665 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 71%|███████▏ | 15/21 [00:07<00:04, 1.43profile/s]2026-08-03 23:01:06,622 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 76%|███████▌ | 16/21 [00:08<00:03, 1.29profile/s]2026-08-03 23:01:07,602 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 81%|████████ | 17/21 [00:09<00:03, 1.19profile/s]2026-08-03 23:01:08,608 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 86%|████████▌ | 18/21 [00:10<00:02, 1.12profile/s]2026-08-03 23:01:09,750 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 90%|█████████ | 19/21 [00:11<00:01, 1.04profile/s]2026-08-03 23:01:10,957 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 95%|█████████▌| 20/21 [00:12<00:01, 1.04s/profile]2026-08-03 23:01:13,623 - INFO - autotuner.py:1699 - flashinfer.jit: [Autotuner]: Skipped 2 unsupported tactic(s) for trtllm::fused_moe::gemm1 (enable debug logs to see details)
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 100%|██████████| 21/21 [00:15<00:00, 1.53s/profile]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm1: 100%|██████████| 21/21 [00:15<00:00, 1.38profile/s]
- (Worker_TP0 pid=165)
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 0%| | 0/21 [00:00<?, ?profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 14%|█▍ | 3/21 [00:00<00:00, 28.11profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 29%|██▊ | 6/21 [00:00<00:01, 8.33profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 38%|███▊ | 8/21 [00:01<00:02, 4.88profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 43%|████▎ | 9/21 [00:01<00:02, 4.11profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 48%|████▊ | 10/21 [00:02<00:03, 3.55profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 52%|█████▏ | 11/21 [00:02<00:03, 3.13profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 57%|█████▋ | 12/21 [00:03<00:03, 2.77profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 62%|██████▏ | 13/21 [00:03<00:03, 2.53profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 67%|██████▋ | 14/21 [00:04<00:02, 2.37profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 71%|███████▏ | 15/21 [00:04<00:02, 2.18profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 76%|███████▌ | 16/21 [00:05<00:02, 2.02profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 81%|████████ | 17/21 [00:05<00:02, 1.87profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 86%|████████▌ | 18/21 [00:06<00:01, 1.74profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 90%|█████████ | 19/21 [00:07<00:01, 1.60profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 95%|█████████▌| 20/21 [00:08<00:00, 1.47profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 100%|██████████| 21/21 [00:09<00:00, 1.17profile/s]
- [AutoTuner]: Tuning trtllm::fused_moe::gemm2: 100%|██████████| 21/21 [00:09<00:00, 2.26profile/s]
- (Worker_TP0 pid=165) 2026-08-03 23:01:23,003 - INFO - autotuner.py:852 - flashinfer.jit: [Autotuner]: Autotuning process ends
- (EngineCore pid=139) INFO 08-03 23:01:24 [shm_broadcast.py:705] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
- (Worker_TP0 pid=165) INFO 08-03 23:01:24 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
- (Worker_TP0 pid=165)
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/5 [00:00<?, ?it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 20%|██ | 1/5 [00:00<00:00, 5.65it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 40%|████ | 2/5 [00:00<00:00, 6.69it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 60%|██████ | 3/5 [00:00<00:00, 6.88it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 80%|████████ | 4/5 [00:00<00:00, 6.28it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 5/5 [00:01<00:00, 3.50it/s]
- Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 5/5 [00:01<00:00, 4.39it/s]
- (Worker_TP0 pid=165)
- Capturing CUDA graphs (decode, FULL): 0%| | 0/4 [00:00<?, ?it/s]
- Capturing CUDA graphs (decode, FULL): 25%|██▌ | 1/4 [00:00<00:00, 4.67it/s]
- Capturing CUDA graphs (decode, FULL): 50%|█████ | 2/4 [00:00<00:00, 5.79it/s]
- Capturing CUDA graphs (decode, FULL): 75%|███████▌ | 3/4 [00:00<00:00, 5.56it/s]
- Capturing CUDA graphs (decode, FULL): 100%|██████████| 4/4 [00:00<00:00, 5.10it/s]
- Capturing CUDA graphs (decode, FULL): 100%|██████████| 4/4 [00:00<00:00, 5.22it/s]
- (Worker_TP0 pid=165) INFO 08-03 23:01:36 [gpu_model_runner.py:6707] Graph capturing finished in 12 secs, took 0.31 GiB
- (Worker_TP0 pid=165) INFO 08-03 23:01:36 [gpu_worker.py:819] CUDA graph pool memory: 0.31 GiB (actual), 0.9 GiB (estimated), difference: 0.59 GiB (187.7%).
- (Worker_TP0 pid=165) INFO 08-03 23:01:36 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
- (EngineCore pid=139) INFO 08-03 23:01:36 [core.py:337] init engine (profile, create kv cache, warmup model) took 180.41 s (compilation: 34.42 s)
- (EngineCore pid=139) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'sliding_attention', 'full_attention'}
- (EngineCore pid=139) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'sliding_attention', 'full_attention'}
- (EngineCore pid=139) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'sliding_attention', 'full_attention'}
- (EngineCore pid=139) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'sliding_attention', 'full_attention'}
- (EngineCore pid=139) [transformers] The tokenizer you are loading from '/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
- (EngineCore pid=139) INFO 08-03 23:01:44 [vllm.py:1042] Asynchronous scheduling is enabled.
- (EngineCore pid=139) INFO 08-03 23:01:44 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native'])
- (EngineCore pid=139) WARNING 08-03 23:01:44 [vllm.py:1486] Auto-initialization of reasoning token IDs failed. Please check whether your reasoning parser has implemented the `reasoning_start_str` and `reasoning_end_str`.
- (EngineCore pid=139) INFO 08-03 23:01:44 [compilation.py:312] Enabled custom fusions: act_quant
- (APIServer pid=48) INFO 08-03 23:01:45 [api_server.py:612] Supported tasks: ['generate']
- (APIServer pid=48) INFO 08-03 23:01:48 [parser_manager.py:37] "auto" tool choice has been enabled.
- (APIServer pid=48) WARNING 08-03 23:01:49 [model.py:1528] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_k': 20, 'top_p': 1.0, 'min_p': 0.0}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}
- (APIServer pid=48) [transformers] The tokenizer you are loading from '/cache/huggingface/models--poolside--Laguna-S-2.1-NVFP4/snapshots/f8fdfcdc4e7b0c474a0102430a8cae0a3a358669' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
- (APIServer pid=48) INFO 08-03 23:01:50 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
- (APIServer pid=48) INFO 08-03 23:01:50 [api_server.py:616] Starting vLLM server on http://0.0.0.0:8888
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:37] Available routes are:
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /docs, Methods: GET, HEAD
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /load, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /version, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /health, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /metrics, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /tokenize, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /detokenize, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/models, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /ping, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /ping, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /invocations, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/responses, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/completions, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/messages, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /generative_scoring, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/completions/render, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
- (APIServer pid=48) INFO 08-03 23:01:50 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
- (APIServer pid=48) INFO: Started server process [48]
- (APIServer pid=48) INFO: Waiting for application startup.
- (APIServer pid=48) INFO: Application startup complete.
- (APIServer pid=48) INFO: 172.19.0.4:34626 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:05:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1259.3 tokens/s, Avg generation throughput: 10.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 0.0%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:05:28 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.1 tokens/s, Avg generation throughput: 18.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 33.0%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:05:38 [loggers.py:273] Engine 000: Avg prompt throughput: 586.6 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 53.6%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:05:48 [loggers.py:273] Engine 000: Avg prompt throughput: 1172.1 tokens/s, Avg generation throughput: 16.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 60.2%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:05:58 [loggers.py:273] Engine 000: Avg prompt throughput: 420.4 tokens/s, Avg generation throughput: 16.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 69.3%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:06:08 [loggers.py:273] Engine 000: Avg prompt throughput: 476.4 tokens/s, Avg generation throughput: 19.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 74.6%
- (APIServer pid=48) INFO 08-03 23:06:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 74.6%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:06:28 [loggers.py:273] Engine 000: Avg prompt throughput: 565.6 tokens/s, Avg generation throughput: 18.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 77.9%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:06:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 13.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:06:48 [loggers.py:273] Engine 000: Avg prompt throughput: 1228.5 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:06:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:07:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:08:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:09:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:10:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:11:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:12:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:12:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:12:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO 08-03 23:12:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 78.8%
- (APIServer pid=48) INFO: 172.19.0.4:34636 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:12:48 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 16.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 79.0%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:12:58 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.7 tokens/s, Avg generation throughput: 18.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 77.4%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:13:08 [loggers.py:273] Engine 000: Avg prompt throughput: 529.7 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 77.6%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:13:18 [loggers.py:273] Engine 000: Avg prompt throughput: 295.8 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.4%, Prefix cache hit rate: 78.5%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:13:28 [loggers.py:273] Engine 000: Avg prompt throughput: 573.3 tokens/s, Avg generation throughput: 20.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 79.0%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:13:38 [loggers.py:273] Engine 000: Avg prompt throughput: 1067.6 tokens/s, Avg generation throughput: 15.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:13:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:13:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:14:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:15:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:16:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:17:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:18:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO 08-03 23:18:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 78.9%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:18:28 [loggers.py:273] Engine 000: Avg prompt throughput: 433.2 tokens/s, Avg generation throughput: 18.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 80.3%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:18:38 [loggers.py:273] Engine 000: Avg prompt throughput: 828.7 tokens/s, Avg generation throughput: 15.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:18:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.1 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO 08-03 23:18:58 [loggers.py:273] Engine 000: Avg prompt throughput: 1666.9 tokens/s, Avg generation throughput: 10.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:19:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:20:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:21:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:22:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:23:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:24:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:25:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO: 172.19.0.4:40398 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:25:18 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 14.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:25:28 [loggers.py:273] Engine 000: Avg prompt throughput: 1046.3 tokens/s, Avg generation throughput: 19.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 80.2%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:25:38 [loggers.py:273] Engine 000: Avg prompt throughput: 586.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 80.2%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:25:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 14.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 79.5%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:25:58 [loggers.py:273] Engine 000: Avg prompt throughput: 1361.4 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 79.5%
- (APIServer pid=48) INFO 08-03 23:26:08 [loggers.py:273] Engine 000: Avg prompt throughput: 785.5 tokens/s, Avg generation throughput: 18.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 79.8%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:26:18 [loggers.py:273] Engine 000: Avg prompt throughput: 892.8 tokens/s, Avg generation throughput: 15.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 80.1%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:26:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 17.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 79.8%
- (APIServer pid=48) INFO 08-03 23:26:38 [loggers.py:273] Engine 000: Avg prompt throughput: 1800.2 tokens/s, Avg generation throughput: 12.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 79.8%
- (APIServer pid=48) INFO 08-03 23:26:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 79.8%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:26:58 [loggers.py:273] Engine 000: Avg prompt throughput: 788.5 tokens/s, Avg generation throughput: 11.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 80.7%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:27:08 [loggers.py:273] Engine 000: Avg prompt throughput: 817.0 tokens/s, Avg generation throughput: 13.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:27:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:27:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:27:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:27:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:27:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:28:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:29:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:30:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:31:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:32:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:33:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:33:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:33:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO: 172.19.0.4:55560 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:33:38 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 15.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:33:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.4%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO 08-03 23:33:58 [loggers.py:273] Engine 000: Avg prompt throughput: 1046.5 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:34:08 [loggers.py:273] Engine 000: Avg prompt throughput: 587.1 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:34:18 [loggers.py:273] Engine 000: Avg prompt throughput: 488.7 tokens/s, Avg generation throughput: 16.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.6%, Prefix cache hit rate: 81.1%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:34:28 [loggers.py:273] Engine 000: Avg prompt throughput: 956.7 tokens/s, Avg generation throughput: 11.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO 08-03 23:34:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 81.0%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:34:48 [loggers.py:273] Engine 000: Avg prompt throughput: 247.1 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:34:58 [loggers.py:273] Engine 000: Avg prompt throughput: 476.7 tokens/s, Avg generation throughput: 19.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:35:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:36:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:36:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:36:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:36:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:36:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:36:58 [loggers.py:273] Engine 000: Avg prompt throughput: 1227.2 tokens/s, Avg generation throughput: 13.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:37:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:38:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:39:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:40:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:41:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO 08-03 23:42:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO: 172.19.0.4:35372 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:43:08 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 17.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.5%, Prefix cache hit rate: 81.4%
- (APIServer pid=48) INFO 08-03 23:43:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1046.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 81.4%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:43:28 [loggers.py:273] Engine 000: Avg prompt throughput: 858.6 tokens/s, Avg generation throughput: 11.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.5%, Prefix cache hit rate: 81.2%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:43:38 [loggers.py:273] Engine 000: Avg prompt throughput: 1171.9 tokens/s, Avg generation throughput: 15.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 80.9%
- (APIServer pid=48) INFO 08-03 23:43:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 80.9%
- (APIServer pid=48) INFO 08-03 23:43:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 80.9%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:44:08 [loggers.py:273] Engine 000: Avg prompt throughput: 297.4 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 81.3%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:44:18 [loggers.py:273] Engine 000: Avg prompt throughput: 731.6 tokens/s, Avg generation throughput: 17.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:44:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:44:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:44:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:44:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO 08-03 23:45:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 81.5%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:45:18 [loggers.py:273] Engine 000: Avg prompt throughput: 735.6 tokens/s, Avg generation throughput: 15.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:45:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:45:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:45:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:45:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:46:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:46:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:46:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:46:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:46:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:46:58 [loggers.py:273] Engine 000: Avg prompt throughput: 756.1 tokens/s, Avg generation throughput: 15.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:47:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:48:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:49:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:50:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:51:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:52:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:60434 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:53:08 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 16.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:53:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.4 tokens/s, Avg generation throughput: 19.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:53:28 [loggers.py:273] Engine 000: Avg prompt throughput: 749.4 tokens/s, Avg generation throughput: 19.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.4%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:53:38 [loggers.py:273] Engine 000: Avg prompt throughput: 610.2 tokens/s, Avg generation throughput: 20.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:53:48 [loggers.py:273] Engine 000: Avg prompt throughput: 477.4 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:53:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:54:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:55:08 [loggers.py:273] Engine 000: Avg prompt throughput: 902.7 tokens/s, Avg generation throughput: 16.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:55:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:55:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:55:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-03 23:55:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:55:58 [loggers.py:273] Engine 000: Avg prompt throughput: 982.3 tokens/s, Avg generation throughput: 15.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-03 23:56:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-03 23:56:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:56:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.8 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-03 23:56:38 [loggers.py:273] Engine 000: Avg prompt throughput: 568.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.7%
- (APIServer pid=48) INFO 08-03 23:56:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.7%
- (APIServer pid=48) INFO 08-03 23:56:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.7%
- (APIServer pid=48) INFO 08-03 23:57:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.7%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-03 23:57:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1559.9 tokens/s, Avg generation throughput: 8.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:57:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:57:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:57:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:57:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:58:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-03 23:59:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:00:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:01:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:48 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:02:58 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:03:08 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:03:18 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:03:28 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:42242 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:03:38 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 15.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:03:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1014.5 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:03:59 [loggers.py:273] Engine 000: Avg prompt throughput: 1046.3 tokens/s, Avg generation throughput: 19.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 82.0%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:04:09 [loggers.py:273] Engine 000: Avg prompt throughput: 1541.9 tokens/s, Avg generation throughput: 13.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.8%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO 08-04 00:04:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.8%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:04:29 [loggers.py:273] Engine 000: Avg prompt throughput: 463.2 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 81.8%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:04:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 16.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO 08-04 00:04:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1563.5 tokens/s, Avg generation throughput: 18.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO 08-04 00:04:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 81.6%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:05:09 [loggers.py:273] Engine 000: Avg prompt throughput: 222.7 tokens/s, Avg generation throughput: 14.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 81.9%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:05:19 [loggers.py:273] Engine 000: Avg prompt throughput: 53.4 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:05:29 [loggers.py:273] Engine 000: Avg prompt throughput: 586.8 tokens/s, Avg generation throughput: 16.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:05:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:05:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:05:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:06:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:07:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:08:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:09:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:10:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:11:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:11:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.5%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO 08-04 00:11:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 17.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:41928 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:11:39 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 18.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.5%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:11:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.7 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:11:59 [loggers.py:273] Engine 000: Avg prompt throughput: 587.1 tokens/s, Avg generation throughput: 12.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 82.1%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:12:09 [loggers.py:273] Engine 000: Avg prompt throughput: 1251.1 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.1%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:12:19 [loggers.py:273] Engine 000: Avg prompt throughput: 545.5 tokens/s, Avg generation throughput: 17.1 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:12:29 [loggers.py:273] Engine 000: Avg prompt throughput: 1134.9 tokens/s, Avg generation throughput: 16.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:12:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:12:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:12:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:13:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:13:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:13:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO 08-04 00:13:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:13:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 16.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 82.0%
- (APIServer pid=48) INFO 08-04 00:13:59 [loggers.py:273] Engine 000: Avg prompt throughput: 1787.8 tokens/s, Avg generation throughput: 12.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 82.0%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:14:09 [loggers.py:273] Engine 000: Avg prompt throughput: 673.9 tokens/s, Avg generation throughput: 15.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:14:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:14:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:14:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:14:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:14:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:15:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:16:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:17:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO 08-04 00:17:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:17:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 10.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.2%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:17:39 [loggers.py:273] Engine 000: Avg prompt throughput: 1403.3 tokens/s, Avg generation throughput: 18.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:17:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:17:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:18:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:19:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:20:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:21:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:22:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:23:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:24:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 19.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.1%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO: 172.19.0.4:49810 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:24:19 [loggers.py:273] Engine 000: Avg prompt throughput: 177.8 tokens/s, Avg generation throughput: 17.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 82.4%
- (APIServer pid=48) INFO 08-04 00:24:29 [loggers.py:273] Engine 000: Avg prompt throughput: 1046.4 tokens/s, Avg generation throughput: 19.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:24:39 [loggers.py:273] Engine 000: Avg prompt throughput: 586.9 tokens/s, Avg generation throughput: 20.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 82.2%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:24:49 [loggers.py:273] Engine 000: Avg prompt throughput: 253.4 tokens/s, Avg generation throughput: 19.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.5%, Prefix cache hit rate: 82.3%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:24:59 [loggers.py:273] Engine 000: Avg prompt throughput: 632.0 tokens/s, Avg generation throughput: 19.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 82.5%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:25:09 [loggers.py:273] Engine 000: Avg prompt throughput: 508.6 tokens/s, Avg generation throughput: 19.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 82.6%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:25:19 [loggers.py:273] Engine 000: Avg prompt throughput: 282.8 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 82.9%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:25:29 [loggers.py:273] Engine 000: Avg prompt throughput: 180.1 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 83.1%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:25:39 [loggers.py:273] Engine 000: Avg prompt throughput: 201.6 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:25:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:25:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:26:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:27:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:28:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:29:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:30:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:30:19 [loggers.py:273] Engine 000: Avg prompt throughput: 1227.3 tokens/s, Avg generation throughput: 11.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:30:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:30:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:30:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:30:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:31:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:32:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:33:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:34:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:35:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.4%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:35:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 18.4 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:35:29 [loggers.py:273] Engine 000: Avg prompt throughput: 778.8 tokens/s, Avg generation throughput: 16.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:35:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:35:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:35:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:36:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:37:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:38:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:39:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:40:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:41:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:41:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.1%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:41:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO 08-04 00:41:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO: 172.19.0.4:35536 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:41:49 [loggers.py:273] Engine 000: Avg prompt throughput: 178.2 tokens/s, Avg generation throughput: 17.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 83.4%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:41:59 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.1 tokens/s, Avg generation throughput: 18.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:09 [loggers.py:273] Engine 000: Avg prompt throughput: 586.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 83.2%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:19 [loggers.py:273] Engine 000: Avg prompt throughput: 1146.3 tokens/s, Avg generation throughput: 15.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.8%, Prefix cache hit rate: 83.1%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:29 [loggers.py:273] Engine 000: Avg prompt throughput: 426.7 tokens/s, Avg generation throughput: 18.9 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 83.2%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:39 [loggers.py:273] Engine 000: Avg prompt throughput: 1420.9 tokens/s, Avg generation throughput: 12.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1077.1 tokens/s, Avg generation throughput: 13.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.2%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:42:59 [loggers.py:273] Engine 000: Avg prompt throughput: 1586.8 tokens/s, Avg generation throughput: 7.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.9%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:43:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:44:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO 08-04 00:44:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:44:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 18.3 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 83.3%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:44:39 [loggers.py:273] Engine 000: Avg prompt throughput: 807.0 tokens/s, Avg generation throughput: 16.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 83.8%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:44:49 [loggers.py:273] Engine 000: Avg prompt throughput: 221.9 tokens/s, Avg generation throughput: 11.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:44:59 [loggers.py:273] Engine 000: Avg prompt throughput: 57.5 tokens/s, Avg generation throughput: 13.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:45:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.5%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:46:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:47:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:48:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.7%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:49:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:50:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:51:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:51:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO 08-04 00:51:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO: 172.19.0.4:39156 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:51:39 [loggers.py:273] Engine 000: Avg prompt throughput: 178.2 tokens/s, Avg generation throughput: 17.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.6%, Prefix cache hit rate: 84.5%
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:51:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1045.1 tokens/s, Avg generation throughput: 18.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.1%, Prefix cache hit rate: 84.3%
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:51:59 [loggers.py:273] Engine 000: Avg prompt throughput: 801.8 tokens/s, Avg generation throughput: 19.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.4%, Prefix cache hit rate: 84.2%
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:52:09 [loggers.py:273] Engine 000: Avg prompt throughput: 519.7 tokens/s, Avg generation throughput: 20.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 84.2%
- (APIServer pid=48) INFO 08-04 00:52:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 24.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 84.2%
- (APIServer pid=48) INFO 08-04 00:52:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 25.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 84.2%
- (APIServer pid=48) INFO 08-04 00:52:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 25.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 84.2%
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:52:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 20.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:52:59 [loggers.py:273] Engine 000: Avg prompt throughput: 1163.1 tokens/s, Avg generation throughput: 18.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:53:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:53:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:53:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO: 172.19.0.4:52812 - "POST /v1/chat/completions HTTP/1.1" 200 OK
- (APIServer pid=48) INFO 08-04 00:53:39 [loggers.py:273] Engine 000: Avg prompt throughput: 1058.4 tokens/s, Avg generation throughput: 13.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:53:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:53:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.7 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:54:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:55:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:56:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:57:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.6 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 22.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:39 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:49 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:58:59 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:59:09 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:59:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 23.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
- (APIServer pid=48) INFO 08-04 00:59:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 21.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.1%
Advertisement
Add Comment
Please, Sign In to add comment