If generative AI can hand you a recipe, agentic AI can be the cook who reads it, checks what's in the fridge, shops for what's missing, and puts dinner on the table. For the many businesses recognizing the force-multiplier effects of agentic AI, the leap from producing answers to executing entire workflows brings a steep new set of demands on enterprise infrastructure. Unlike chatbots, which typically map a single user message to a single inference call, agentic tasks can spawn hundreds of calls—or more!—running in sequence and in parallel, each requiring seamless coordination among GPU compute, CPU compute and orchestration, and tool execution. For IT decision-makers planning strategic AI infrastructure investments, those resource-intensive demands raise important questions: which platform, and which configuration of that platform, gives them the best chance of meeting demand as workloads grow? For any business embracing agentic AI—especially those that also need to keep sensitive data under local control—choosing the right server is a decision with long-term consequences. 

The Dell PowerEdge XE7740, powered by Intel Xeon 6747P processors and NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, promises to meet agentic demands while offering the security and cost control of an on-premises solution. But server specifications alone won't help IT buyers understand how many agents a platform can actually support, what performance users will experience, or what the solution will cost to run over time. Answering those questions calls for real-world, fact-based sizing data. 

To provide that data, we tested multiple GPU configurations of the PowerEdge XE7740 on two realistic agentic AI workloads—one for financial services and one for manufacturing—using a custom benchmark and framework that runs full-stack, complete end-to-end agent workflows. We configured the server with two Intel Xeon 6747P CPUs and scaled from four to six to eight NVIDIA GPUs to measure how performance changed. We also examined the tokenomics of the 8-GPU configuration, comparing its cost per million tokens against Amazon Bedrock running the same model. 

We found that performance scaled clearly as we added GPUs. With the financial services workload, the 8-GPU configuration supported up to 74 simultaneous AI agents while holding task-completion latency under 30 seconds, and it delivered up to 1,179 tokens per second. With the manufacturing workload, doubling the GPU count nearly doubled both supported agents and throughput. The cost advantage was just as striking. At full utilization, the PowerEdge XE7740 achieved a cost as low as $0.09 per million tokens—just 5.4 percent the cost of the comparable Amazon Bedrock solution over five years. 

Our analysis suggests that organizations planning agentic AI deployments can meet real-world demand, maintain control over their data, and reduce long-term costs with the Dell PowerEdge XE7740, choosing the four- or eight-GPU configuration that best fits their workloads.  

To learn more about our Dell PowerEdge XE7740 agentic AI performance and tokenomics findings, check out the report below.