Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A custom PC with an RTX 3090 was reported to generate local AI output at 140 tokens per second, compared with 85 tokens per second on a 64GB M5 Max Mac Studio. That is a notable result for the PC on sustained generation, but it is one reported comparison—not proof that every RTX 3090 setup is faster, or that the Mac is slower for every local-AI workload.

What the comparison reports

Geeky Gadgets’ September 29, 2026 summary of a comparison by The Stack names Qwen 3.6, described in the article as a 35-billion-parameter model. It reports these results for the tested configurations:

Measure RTX 3090 custom PC 64GB M5 Max Mac Studio
Sustained generated output 140 tokens per second, as reported by Geeky Gadgets 85 tokens per second, as reported by Geeky Gadgets
Short-prompt processing About 3,000 tokens per second, as reported by Geeky Gadgets About 3,000 tokens per second, as reported by Geeky Gadgets
Longer-prompt processing Not stated in Geeky Gadgets’ accessible text About 2,000 tokens per second, as reported by Geeky Gadgets
Context at full speed About 90,000–150,000 tokens, as reported by Geeky Gadgets 262,000 tokens, as reported by Geeky Gadgets
System price Approximately $2,000 for the article’s custom PC estimate; not a current quote $3,799 for the article’s Mac Studio estimate; not a current quote

All benchmark and price figures in this table are Geeky Gadgets’ account of the comparison, not independently reproduced measurements or verified current prices. The summary does not establish that the systems used identical model files or disclose quantization, runtime versions, batch size, prompt and output lengths, or test conditions. The linked The Stack video page does not provide accessible methodology details sufficient to resolve those questions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What 140 versus 85 tokens per second means

Tokens per second can refer to different stages of an AI request. Prompt processing measures how quickly the system reads the supplied text; generated-token throughput measures how quickly it produces the response. The reported short-prompt figure of about 3,000 tokens per second is not directly comparable to the 140- and 85-token-per-second sustained output figures. A fast prompt-processing rate does not necessarily mean the answer will be generated at the same rate.

#1 Best Overall
Sale
NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)
  • Item Package Dimension - 15.0L x 12.25W x 4.25H inches
  • Item Package Weight - 6.0 Pounds
  • Item Package Quantity - 1
  • Product Type - VIDEO CARD

For a workload that spends most of its time generating long responses, the reported 140-versus-85 result favors the RTX 3090 configuration. It should not be treated as a guaranteed ratio across models or software: context length, model implementation, quantization, and runtime can all affect throughput, and the comparison summary does not document enough details to isolate their effects.

Context capacity changes the tradeoff

The reported comparison gives the Mac a larger context window at full speed: 262,000 tokens versus roughly 90,000–150,000 for the RTX 3090 setup. Context is the material a model can consider in a request, including the prompt and conversation history; a larger context can matter for tasks involving long documents or extended conversations. The figures describe the tested systems and should not be assumed for every model, quantization, or runtime.

The article also says 32GB of PC system RAM was generally sufficient for its described Qwen workload, and that 128GB added little for that setup. That is a workload-specific observation, not a general RAM recommendation for local AI.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Memory figures are not a like-for-like comparison

The Mac’s 64GB refers to unified memory. Apple lists M5 Max Mac Studio configurations with 48GB, 64GB, or 128GB of unified memory in its Mac Studio technical specifications. A discrete RTX 3090 uses its own GPU memory alongside the PC’s system RAM; the comparison article’s linked video title identifies a 24GB card, but the accessible page does not provide further test details. These memory pools are not interchangeable, so comparing “64GB” with “24GB” alone does not tell you which system will run a particular model or context more effectively.

Quick Recap

SaleBestseller No. 1
NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)
NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)
Item Package Dimension - 15.0L x 12.25W x 4.25H inches; Item Package Weight - 6.0 Pounds; Item Package Quantity - 1
$1,864.99
SaleBestseller No. 3
MSI Gaming GeForce RTX 3090 24GB GDRR6X 384-Bit HDMI/DP Nvlink Torx Fan 3 Ampere Architecture OC Graphics Card (RTX 3090 VENTUS 3X 24G OC) (Renewed)
MSI Gaming GeForce RTX 3090 24GB GDRR6X 384-Bit HDMI/DP Nvlink Torx Fan 3 Ampere Architecture OC Graphics Card (RTX 3090 VENTUS 3X 24G OC) (Renewed)
Digital Maximum Resolution - 7680 X 4320; Output- Displayport X 3 (V1.4A) / Hdmi 2.1 X 1; Memory Interface- 384-Bit
$1,849.99
Bestseller No. 5
Best Value
ASUS ROG Strix NVIDIA GeForce RTX 3090 Gaming Graphics Card- PCIe 4.0, 24GB GDDR6X, HDMI 2.1, DisplayPort 1.4a, Axial-tech Fan Design, 2.9-Slot
  • Memory Speed:19.5 Gbps.Digital Max Resolution:7680 x 4320
  • NVIDIA Ampere Streaming Multiprocessors: The building blocks for the world’s fastest, most efficient GPU, the all-new Ampere SM brings 2X the FP32 throughput and improved power efficiency.
  • 2nd Generation RT Cores: Experience 2X the throughput of 1st gen RT Cores, plus concurrent RT and shading for a whole new level of ray tracing performance.
  • 3rd Generation Tensor Cores: Get up to 2X the throughput with structural sparsity and advanced AI algorithms such as DLSS. Now with support for up to 8K resolution, these cores deliver a massive boost in game performance and all-new AI capabilitiesAvoid using unofficial software
  • Axial-Tech Fan Design has been newly tuned with a reversed central fan direction for less turbulence.
Rank #3
Sale
MSI Gaming GeForce RTX 3090 24GB GDRR6X 384-Bit HDMI/DP Nvlink Torx Fan 3 Ampere Architecture OC Graphics Card (RTX 3090 VENTUS 3X 24G OC) (Renewed)
  • Digital Maximum Resolution - 7680 X 4320
  • Output- Displayport X 3 (V1.4A) / Hdmi 2.1 X 1
  • Memory Interface- 384-Bit
  • Package Quantity-1

Which setup fits your priorities?

  • Choose the RTX 3090 PC if: sustained generation speed is the priority and you are comfortable selecting and configuring a custom system. The reported result favors this configuration for generated output, but confirm performance with the model and runtime you intend to use.
  • Consider the Mac Studio if: the larger context capacity reported in this comparison matters more to your workload, or you value a ready-to-use integrated system. Apple’s specifications confirm several M5 Max memory configurations; they do not validate the benchmark result.
  • Compare complete, current costs: the article’s approximately $2,000 PC and $3,799 Mac figures are dated estimates, not verified quotes. A PC total depends on its components and the RTX 3090’s condition; assess the complete configuration rather than comparing only the graphics card with the Mac’s purchase price.

How to use the result before buying

  1. Match the workload. Decide whether you need faster response generation, long-context document handling, or both.
  2. Verify the exact configuration. Check the model, quantization, runtime, and memory available to the system you plan to use; the summary does not disclose enough detail to guarantee the reported results will transfer.
  3. Seek workload-specific measurements. Compare prompt-processing and generated-token throughput separately, at context lengths similar to your own tasks.
  4. Price the whole system at purchase time. Treat the article’s estimates as historical context, and account for the condition and price of the GPU and the rest of the PC build.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.