Sobes.tech
Senior

What other optimization methods can be used to speed up inference or at least fit the model into 12 GB VRAM RTX 3090?