News

The developer deployed the Kimi K3 model on 8GB of memory and ran it successfully, but... the speed was 33 seconds/Token

1 min read
Source: landian.news
#artificial intelligence A developer deployed the Kimi K3 model locally, which only requires 8.24GB of memory to run, but the running speed is 33 seconds/Token (note the unit). If the memory is increased to 128GB, 20 seconds/Token can be achieved, but this is basically the upper limit of performance. Continuing to heap memory cannot continue to improve efficiency. Currently, developers have made the inference engine developed based on C99 open source, and interested users can study it. View the full text: https://ourl.co/114176