(IIUC, -ngl [NUM_LAYERS] specify number of layers to offload to GPU in llama.cpp, 999 on most use cases might as well be -1)