Andromeda Cluster - Solutions Architect
Requirements
• ● 5+ years in customer-facing technical roles, including 2+ years in pre-sales (Solutions Engineer, Sales Engineer, Solutions Architect, or specialist SA). • ● Direct experience selling or supporting GPU compute at a neocloud or GPU provider, or as an AI/HPC specialist at a hyperscaler or NVIDIA. • ● Demonstrated ability to take a customer from evaluation into production and stay accountable for the outcome. • ● Real fluency in large-scale training and inference: distributed training frameworks, multi-node topologies, InfiniBand/RoCE, storage and checkpointing, and where these break at scale. • ● Comfort with Kubernetes and SLURM as scheduling environments customers actually run in. • ● Enough Python to build a benchmark, a prototype, or an API integration yourself rather than waiting on engineering. • ● A track record of owning technical evaluations in complex, multi-stakeholder deals and changing the outcome. • ● Exceptional communication, able to hold a deep conversation with a distributed systems engineer and a CFO. • Strong Candidates May Have • ● Experience with frontier labs or AI-native companies as customers. • ● Performance benchmarking, MFU analysis, or total-cost-of-training modeling. • ● Having been the first or earliest SE somewhere. • What Success Looks Like • ● POCs convert reliably because evaluations are well-scoped, technically credible, and tied to customer ROI. • ● Customers reach their first production training run fast, and what they saw in the POC is what they get. • ● Accounts grow because you saw the constraint coming before the customer raised it. • ● The assets you build let the next five SAs ramp in weeks instead of quarters.
Responsibilities
• Win the technical evaluation • ● Partner with Sales to qualify opportunities, lead technical discovery, and scope evaluations against the customer's real success criteria — not a generic checklist. • ● Design cluster and workload architectures that map to business outcomes: time-to-first-token, tokens/sec, MFU, cost per training run, reliability targets. • ● Own POCs end-to-end — scope, benchmarks, success metrics, timeline, stakeholder alignment. • Land it and grow it • Land it and grow it • ● Take customers from signature to first successful production training run: provisioning, environment setup, validation benchmarks, and the unglamorous debugging in between. • ● Serve as the named technical owner for your accounts after launch. You’re responsible for architecture reviews, capacity planning, performance and cost optimization, and technical workload reviews. • ● Spot expansion before the customer asks: where they're capacity-constrained, what's coming on their roadmap, and what it will take to serve it. • Build the function • Build the function • ● Create the assets the SA team runs on: demo environments, benchmarking harnesses, reference architectures, onboarding runbooks, evaluation playbooks, and competitive material. • ● Be the highest-signal feedback loop into Product, Engineering, and Research — recurring gaps, competitive losses, and what customers actually ask for once they're in production.
Benefits
• ● High-growth environment: Get in early at a company at the center of the AI infrastructure boom • High-growth environment: • ● Ownership: First HPC Architect for the solutions engineering team, you’ll get to build this function from the ground up • Ownership: • ● Competitive compensation: + meaningful equity • ● Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT