Trusted Partners
Start from the job rather than the parts list. Fine-tuning on a single card, serving a model to a small team, running vision pipelines over recorded video, and preparing datasets that never reach a GPU at all are four different machines, and only one of them is mostly a question of graphics. Write down what has to run overnight, what has to stay interactive, and how much data must sit on local storage at once. The specification falls out of those three answers.
Two rules then survive nearly every workload. Have enough system memory that the data pipeline is never the reason a card is idle, and enough storage throughput that the pipeline is never the reason the memory is empty. Get either wrong and a more expensive card simply spends more money waiting.