NovaSynth by Noveum

NovaSynth by Noveum

17/09/2026
Sponsored Link

NovaSynth by Noveum Investment Report

Category: AI-agent evaluation, observability, and synthetic voice testing

Company Stage: Early-stage commercial SaaS

Founder or Founders: Shashank Agarwal and Additi Upadhyay

Headquarters: San Francisco, California, United States

Funding: Not publicly disclosed; no reliable announcement of an institutional Noveum financing was found

Business Model: Organization-level SaaS with usage credits and custom enterprise deployments

Product Hunt Launch Date: September 17, 2026

Report Date: September 20, 2026

Investment MetricAssessment
Venture Potential73/100
Unicorn PathConditional
Valuation AttractivenessNot Assessable
Evidence Confidence58/100
Final DecisionDD

Executive Summary

NovaSynth is Noveum’s pre-production testing system for voice and text agents. It generates synthetic callers with configurable goals, accents, moods, interruptions, background noise, network conditions, and adversarial behavior, then connects them to an agent over real telephone, LiveKit, or HTTP channels. Calls are traced and evaluated using audio, transcript, safety, and business-outcome scorers (NovaSynth; documentation).

The narrow initial buyer is an engineering or quality-assurance team operating customer-facing voice agents in areas such as call centers, collections, financial services, healthcare, insurance, travel, and sales. These customers have meaningful willingness to pay because manual call testing is difficult to reproduce, while production failures can create direct financial, compliance, and reputational costs.

The strongest investment signal is that NovaSynth is part of a broader evaluation stack rather than a single-purpose testing utility. Noveum combines open-source tracing, production evaluation, synthetic testing, root-cause analysis, and proposed fixes. It reports production deployments with MyOperator and DarGlobal, publishes transparent self-serve pricing, supports on-premise deployment, and is hiring technical staff (Noveum platform; careers).

The most important concern is insufficient independently verified commercial traction. Noveum has not disclosed revenue, paying customers, growth, retention, gross margin, or funding. Its performance evidence is largely company-published, including a MyOperator case study. The open-source tracing SDK has visible technical activity but limited external adoption, with 14 GitHub stars at the time reviewed (GitHub).

The market and founding team are strong enough to justify formal diligence. The decision is DD, not Invest, because valuation, financing terms, customer retention, unit economics, and the reliability of Noveum’s evaluation methodology remain unknown.

Product Overview

Voice agents can pass scripted transcript tests while failing under interruptions, accents, noisy environments, ambiguous answers, silence, poor connections, or multi-turn behavioral changes. Human testers struggle to reproduce these conditions consistently.

NovaSynth creates a persona and scenario, connects a synthetic user to the target agent, runs individual or batch sessions, captures the resulting trace and audio, and evaluates the interaction. It supports phone connections, LiveKit voice or text sessions, and HTTP chat. Direct Vapi, Retell, ElevenLabs Conversational, Pipecat, and arbitrary WebSocket endpoints are not currently executed natively; some can still be tested through a phone number or compatible HTTP endpoint (connection documentation).

Features include batch persona-scenario matrices, up to 2,000 concurrent lines for arranged load tests, tool virtualization, multilingual personas, regression testing, 30-plus voice metrics, and more than 100 evaluation scorers across Noveum’s platform. Its documentation notes that NovaSynth has no built-in nightly scheduler; customers must use the public API for scheduled testing.

Pricing is organization-based rather than per seat. Plans range from free to $69, $99, $199, and $599 per month, plus custom Enterprise contracts. The free plan includes 2,500 monthly credits, approximately 100 voice minutes. Voice testing consumes 25 credits per minute; add-on pricing implies roughly $0.125–$0.20 per voice minute depending on the credit purchase method (pricing).

The product replaces manual QA calls, spreadsheet test plans, ad hoc prompt experiments, and transcript-only evaluation.

Product Quality: Technically substantial and well documented, with useful voice-specific testing capabilities; reliability and customer value at production scale need independent validation.

Founder and Team Assessment

Shashank Agarwal is Noveum’s founder and CEO. His professional profile reports prior engineering leadership at Activeloop and Levity, work on the original AWS SageMaker launch team, and machine-learning engineering at Amazon Prime Video. These experiences are directly relevant to observability, scalable infrastructure, and model operations, although performance figures attributed to his prior work are self-reported (Agarwal profile).

Additi Upadhyay is co-founder and chief growth officer. Her profile reports prior testing experience at IBM, product work at Playo and Spenmo, and product and growth work at SambaNova. This creates credible founder-market fit across QA, enterprise AI, product, and go-to-market functions (Upadhyay profile).

Both founders also identify themselves as active founders of API.Market. That may provide distribution and infrastructure advantages, but it creates a focus and ownership question: investors must determine how management time, employees, intellectual property, revenue, and expenses are divided between the two businesses.

LinkedIn lists Noveum as a 2–10-person company with ten associated profiles. The careers page showed openings for a senior AI engineer and a full-stack engineer, supporting a current hiring signal. Exact full-time headcount and founder ownership are not verified.

Founder Assessment: Strong technical and product-growth backgrounds, offset by key-person concentration and possible divided focus across Noveum and API.Market.

Market Opportunity

The initial segment is companies running customer-facing voice agents at sufficient volume that manual QA and production incident review become expensive. This is narrower than the total contact-center or AI software market.

A bottom-up scenario can be framed as follows:

  • 10,000–30,000 organizations globally operating or developing production voice and complex conversational agents;
  • $7,000–$25,000 potential annual contract value across testing, tracing, evaluation, and remediation;
  • 10% obtainable penetration over time.

This implies approximately $7M–$75M in ARR. These are analyst assumptions, not verified market counts. NovaSynth alone may support a meaningful SaaS business, but the broader Noveum platform is more likely to support venture-scale revenue through continuous production observability, enterprise deployment, and usage expansion.

Adjacent opportunities include text-agent testing, multi-agent workflow evaluation, release gating, compliance evidence, automated remediation, agent-security testing, and on-premise deployments. Geographic expansion is possible because the setup documentation supports numerous languages, although accurate multilingual testing must be validated rather than inferred from a language selector.

Market timing is favorable because more enterprises are deploying nondeterministic agents into customer-facing workflows. However, the category is attracting numerous well-funded and technically capable competitors.

Traction and Growth Signals

NovaSynth ranked fifth on Product Hunt on September 17, 2026 (leaderboard). That is a useful launch signal but provides little evidence of recurring commercial demand.

Noveum names MyOperator and DarGlobal as production customers. Its self-published MyOperator case study claims a four-line integration, 20%–30% lower AI-engineering overhead, a 200-fold faster workflow, and improvement in the evaluated success rate from 84% to above 95% (case study). These results are promising but were not independently audited and concern the broader Noveum platform, not necessarily standalone NovaSynth usage.

The company claims its infrastructure processes more than six million traces per day and has been load-tested at 60 million spans per day. These are company-reported capacity and usage figures; the split between production customers, internal testing, and free users is not disclosed.

The Apache-2.0 tracing SDK supports Python, TypeScript, LangChain, LangGraph, LiveKit, Pipecat, and CrewAI. Public GitHub activity confirms a functioning developer product, but low repository stars suggest that open-source-led distribution is still early.

The most important missing metrics are ARR, paid organizations, NovaSynth test volume, customer concentration, conversion from free to paid, logo retention, net revenue retention, and post-launch pipeline.

Traction Assessment: Credible product deployment signals and one detailed customer case study, but commercially unverified.

Competitive Position

Direct voice-testing competitors include Hamming AI, Cekura, and Coval. Broader evaluation and observability competitors include Braintrust, Langfuse, Arize Phoenix, Confident AI, and AgentOps.

Free and manual alternatives include internal QA teams, scripted telephone calls, custom evaluation harnesses, OpenTelemetry instrumentation, and open-source Langfuse or Phoenix deployments. Voice-platform vendors such as Vapi, Retell, ElevenLabs, and LiveKit could also bundle deeper testing and monitoring into their core platforms.

Noveum differentiates through an integrated loop: tracing production failures, evaluating them, generating synthetic scenarios, testing candidate fixes, and returning proposed code or prompt changes. NovaSynth’s audio-level scoring and repeatable caller behavior are more specialized than generic LLM evaluation.

Switching costs could become meaningful once a customer has instrumented production agents, created datasets and scorers, and embedded Noveum into release gates. However, the open SDK and standard OpenTelemetry orientation also make data more portable. Proprietary advantage would need to come from scoring accuracy, accumulated failure data, integrations, and demonstrated remediation quality.

If the largest platform launched the same feature within six months, customers would continue using Noveum only if its scorers produced demonstrably better results, its simulations found failures competitors missed, or its closed-loop fixes generated measurable engineering savings. Public evidence is not yet sufficient to establish that advantage.

Defensibility Assessment: Medium

Business Model and Economics

Noveum combines recurring plans, metered credits, tracing overages, storage charges, and custom enterprise deployments. Public annualized subscription values range from approximately $690 for Pro to $7,188 for Scale, before add-ons and enterprise contracts.

NovaSynth’s variable costs include telephony or SIP capacity, speech recognition, speech synthesis, LLM inference for synthetic callers and scorers, audio storage, and cloud execution. At $0.125–$0.20 of incremental revenue per voice minute, gross margin could be sensitive to model and telephony choices. Actual cost per minute and gross margin are not disclosed.

The broader platform provides better economics than voice simulation alone: rule-based scorers have low marginal cost, while tracing, storage, evaluation, remediation, support, and private deployment create expansion revenue. Conversely, on-premise installations and complex enterprise integrations may require substantial support and professional services.

There are no App Store fees. Standard payment-processing fees apply to self-serve subscriptions. Enterprise sales cycles and security reviews may create material acquisition costs.

Unicorn Path

Assuming an 8× ARR multiple for a scaled infrastructure SaaS company with durable growth, meaningful enterprise revenue, and acceptable gross margins:

Required ARR = $1 billion ÷ 8 = approximately $125 million.

At the current $599 monthly Scale price, Noveum would need approximately 17,400 equivalent Scale customers. At a hypothetical $25,000 enterprise ACV, it would require 5,000 enterprise customers. At $100,000 ACV, it would require 1,250 customers. The enterprise ACVs are analyst scenarios because contract values are not disclosed.

NovaSynth cannot credibly reach this scale as an isolated testing module at current self-serve prices. A unicorn outcome requires Noveum to become a broad system of record for agent reliability: production tracing, evaluation, simulation, policy enforcement, automated remediation, and auditable release gates. It would also require repeatable enterprise distribution, international reach, strong net retention, and gross margins that remain attractive as voice usage expands.

Unicorn Path: Conditional

Valuation Assessment

No reliable public information was found regarding Noveum’s total funding, institutional investors, latest round, current fundraising status, valuation, SAFE cap, or secondary transactions. Third-party databases describing Noveum as unfunded and estimating revenue are not sufficiently reliable to treat as verified facts.

No responsible private valuation range can be derived from public product quality, Product Hunt ranking, or company-reported usage alone. Financing and acquisition comparables in AI observability vary widely based on ARR, growth, retention, customer concentration, and strategic value.

Valuation Attractiveness: Not Assessable

Required information includes current ARR, monthly growth, gross margin by product, net retention, customer concentration, burn, runway, cap table, financing amount, post-money valuation, option pool, and liquidation preferences.

Key Risks

  1. Commercial opacity: Revenue, paid customer count, growth, and retention are undisclosed.
  2. Crowded category: Multiple voice-testing and AI-observability competitors offer overlapping features.
  3. Evaluation reliability: LLM-based judges can produce inconsistent or biased scores.
  4. Usage economics: Voice simulation combines telephony, STT, TTS, LLM, and storage costs.
  5. Founder focus: Both founders also operate API.Market.
  6. Customer concentration: Only two named production customers were readily verifiable.
  7. Security readiness: SOC 2 Type II is described as in progress rather than completed.
  8. Platform integration gaps: Direct execution for several prominent voice platforms is not yet supported.
  9. Automated-fix liability: Incorrect prompt or code recommendations could degrade production agents.
  10. Low current switching barriers: Competing OpenTelemetry-compatible platforms can ingest similar traces.

Final Assessment

Venture Potential: 73/100

CategoryScore
Market Size and Expansion Potential17/20
Traction and Growth Evidence11/20
Founder and Team13/15
Product Strength9/10
Distribution Potential9/15
Business Model and Economics8/10
Defensibility6/10
Total73/100

The strongest elements are founder-market fit, technical product breadth, enterprise deployment options, and expansion from testing into continuous agent reliability. The weakest are limited verified traction, crowded competition, and unknown unit economics.

Evidence Confidence: 58/100

Product capabilities, documentation, pricing, founders, GitHub activity, hiring, and two customer references are verifiable. Throughput and performance outcomes are company-reported. Market sizing and enterprise ACVs are analyst assumptions. Financial results, retention, funding, valuation, and cap-table information remain unavailable.

Final Decision: DD

Noveum warrants a founder meeting and data-room review. The company appears technically capable and addresses a potentially large infrastructure need, but the public record does not establish commercial momentum, defensibility, or an attractive investment price.

Upgrade Conditions

  • Verified ARR above $1M with sustained growth.
  • At least 20 paying mid-market or enterprise customers with low concentration.
  • More than 70% six-month customer retention and strong usage expansion.
  • Gross margin above 70% after telephony and AI-inference costs.
  • Independent customer references confirming reduced incidents or QA expense.
  • Completed SOC 2 Type II certification.
  • Evidence that NovaSynth materially outperforms competing testing platforms.
  • Clear founder focus and separation from API.Market.

Downgrade Conditions

  • Weak free-to-paid conversion or low testing frequency after onboarding.
  • Gross margin deterioration as voice volume increases.
  • Loss of named enterprise customers.
  • Larger platforms bundle equivalent simulation and evaluation.
  • Evaluation scores prove poorly correlated with production outcomes.
  • Material security, privacy, or automated-remediation incidents.
  • Product activity declines or founders prioritize another company.

Questions for Further Diligence

  1. What are current ARR, MRR, and monthly revenue growth?
  2. How many paying organizations use NovaSynth versus other Noveum modules?
  3. What are 30-, 90-, and 180-day logo and usage retention?
  4. What are gross and net revenue retention?
  5. What is gross margin per voice-test minute after telephony, STT, TTS, inference, and storage?
  6. How much of reported trace volume comes from paying external customers?
  7. What are average enterprise ACV, sales-cycle length, CAC, and payback period?
  8. How are scorer accuracy and synthetic-caller realism independently benchmarked?
  9. What percentage of proposed NovaPilot fixes are accepted and retained in production?
  10. How are staff, intellectual property, and founder time divided between Noveum and API.Market?
  11. What are current burn, cash balance, runway, and team structure?
  12. What are the cap table, current valuation, proposed round terms, and liquidation preferences?

Sources