Trusted Partners
The arithmetic is simple enough to do before you shop. At full sixteen-bit precision a model needs roughly two gigabytes of memory per billion parameters. Four-bit quantisation, the format most local users settle on, brings that down to around half a gigabyte per billion plus a modest overhead. An eight-billion-parameter model therefore lands near five gigabytes, a fourteen-billion model near nine, and a seventy-billion model near forty.
Then add the context. The key-value cache holding a conversation grows with its length and with how many requests run at once, and it occupies the same memory as the weights. A card that fits a model with two gigabytes to spare will manage a short exchange and run out partway through a long document. Leaving headroom for context is what separates a setup that works in a demonstration from one that works all day.