A voice agent has only a fraction of a second to understand a caller, retrieve the right context, make a decision and begin responding. It may also handle account information, payment details, health data or other sensitive information while doing so.
That makes model selection only one part of the architecture.
Enterprises must also decide where inference runs, which systems receive conversation data, how the model accesses business context, what controls govern its actions and how the complete interaction is recorded.
For high-volume, bounded voice workflows, a self-hosted small language model can provide a compelling answer. Its smaller computational footprint makes customer-controlled deployment more practical while preserving the speed required for live conversations.
It is not the right choice for every workload. It is valuable when model capability, infrastructure control and production economics align with the task.
What makes a language model “small”?
A small language model, or SLM, is a language model with a smaller parameter count and computational footprint than a large general-purpose model. It requires less memory and processing capacity, making it easier to deploy in resource-constrained or privately controlled environments.
Size and specialization are separate characteristics.
An SLM may begin as a general-purpose model and then be adapted for a particular domain or set of tasks. That adaptation can involve fine-tuning, retrieval, structured instructions, deterministic policies or a combination of these techniques.
For enterprise voice AI, the objective is not simply to use a smaller model. The objective is to use the smallest model that can reliably meet the workload’s requirements for comprehension, reasoning, structured output, policy adherence and tool use.
Why voice AI creates a different architectural problem
Voice is synchronous. The caller experiences every delay.
A production voice system must coordinate several components:
Caller → telephony and media streaming → voice activity detection → speech recognition → context and policy retrieval → language model → tools and enterprise systems → speech synthesis → caller
Each stage affects the experience. Model inference is important, but it is not the only source of latency.
Queueing, long prompts, retrieval, tool execution, speech recognition and speech synthesis may all add delay. A useful latency measurement therefore includes more than model response time. It should distinguish:
- Time to first text
- Time to first audio
- Speech-recognition latency
- Model time to first token
- Tool-execution time
- Text-to-speech latency
- End-to-end response time
- Performance at median, p95 and p99
- Performance under expected concurrent-call volume
This matters because self-hosting does not automatically make a voice agent faster. A well-operated, appropriately located deployment can reduce network dependency and improve control over the inference path. An undersized deployment can create queueing and perform worse than a managed service.
NVIDIA’s inference guidance shows that time to first token includes network, prefill and queueing time, and that latency increases when concurrency exceeds available serving capacity.
The relevant question is therefore not “Is the model self-hosted?” It is “Can the complete voice system meet its latency objective at production load?”
What self-hosting actually means
The term “self-hosted” is often applied too broadly. Four common deployment models should be distinguished.
Managed model API
The model runs in provider infrastructure and is operated by the provider. Customer control depends on the provider’s service controls and contract.
Dedicated managed environment
The model runs in provider infrastructure reserved for one customer and is operated by the provider. This offers greater isolation while the provider retains operational control.
Customer VPC
The model runs in the customer’s cloud account and network boundary. The customer, vendor or both may operate it. This gives the customer direct control over networking, access, keys and surrounding services.
On-premises or private cloud
The model runs in customer-controlled infrastructure and is operated by the customer, vendor or both. This offers maximum infrastructure control and operational responsibility.
A dedicated single-tenant service can provide meaningful isolation, but it is different from deploying the model inside a customer-owned cloud account.
That distinction affects who controls:
- Administrative access
- Encryption keys
- Network egress
- Model artifacts
- Logs and retention
- Software updates
- Runtime dependencies
- Capacity and failover
- Security monitoring
- Incident response
A serious deployment decision must define each of these controls explicitly.
What a customer-hosted SLM can provide
A clearer data boundary
When the model runs inside a customer-controlled VPC or private environment, prompts and model outputs can remain within that environment.
The surrounding voice stack must follow the same boundary. If audio, transcripts, retrieved documents or generated speech are sent to external ASR, TTS, analytics or observability services, the conversation is not fully contained simply because the language model is local.
The architecture should document every point where data is processed, stored or transmitted.
Direct infrastructure control
A customer-hosted deployment can give the enterprise direct control over:
- Private network connectivity
- Identity and role-based access
- Customer-managed encryption keys
- Data-residency configuration
- Logging and retention
- Internet egress
- Model versioning
- Upgrade and rollback schedules
Managed AI services can also provide strong enterprise controls. Azure, Amazon Bedrock and Google Cloud offer combinations of private networking, customer-managed encryption, regional processing and restrictions on the use of customer data.
Customer hosting becomes valuable when the enterprise requires more direct control than those managed-service arrangements provide.
Greater control over performance engineering
Running the inference environment gives the enterprise and model provider control over:
- Hardware selection
- Quantization
- Inference runtime
- Batching
- Maximum concurrency
- Replica count
- Geographic placement
- Autoscaling behaviour
- Admission control
- Model and prompt versions
A smaller model can make these choices easier to implement within a practical GPU and cost envelope.
Performance still has to be measured against real production traffic. A benchmark without hardware, concurrency, input length, output length and percentile measurements is not enough to support a latency claim.
Capacity-based economics
Managed model APIs commonly charge according to tokens, requests or usage. Customer-hosted inference shifts more of the cost toward provisioned infrastructure and operations.
At sustained volume, this can make costs more predictable. At low or highly variable utilization, managed inference may remain more economical.
The comparison should include:
- GPU and infrastructure cost
- Expected utilization
- Peak concurrent calls
- Redundant capacity
- Engineering and support
- Monitoring and security operations
- Cost per completed conversation
- Cost of model upgrades
- Managed-service pricing for the same workload
A smaller model improves the economics only when it satisfies the required quality threshold.
Self-hosting does not create security or auditability by itself
A private deployment can strengthen control. It can also transfer substantial responsibility to the enterprise.
The deployment still needs:
- Signed and verified model artifacts
- Secrets management
- Runtime and driver patching
- Network segmentation
- Least-privilege access
- Vulnerability monitoring
- Backup and disaster recovery
- High-availability design
- Capacity monitoring
- Incident-response procedures
- Controlled model promotion and rollback
The same applies to auditability. Owning the infrastructure does not automatically explain what happened during a customer interaction.
A useful execution trace should capture:
- The caller’s intent
- The context retrieved for that turn
- The applicable policy or decision rule
- The model and prompt version
- The proposed tool action
- Any authorization or human approval
- The tool result
- The resulting state change
- The final system-of-record outcome
- Any escalation or exception
This produces traceability. It shows which inputs, rules and actions contributed to the outcome without claiming that the internal reasoning of a neural model can always be explained.
When a self-hosted SLM makes sense
A customer-hosted SLM is especially relevant when the workload has:
- High and relatively predictable interaction volume
- Tight time-to-first-audio requirements
- Repetitive or bounded customer intents
- Strict data-boundary requirements
- A need to pin model versions
- Restrictions on external model dependencies
- Stable domain terminology and workflows
- Sufficient infrastructure and operational support
- A quality benchmark that the smaller model can consistently pass
Examples can include appointment management, order servicing, account support, collections, claims intake, employee assistance and other structured voice workflows.
When another architecture may be better
A managed model may remain the better choice when:
- Traffic is small or highly variable
- The team is still validating the use case
- The workload requires frontier-level reasoning
- The organisation lacks model-serving expertise
- Rapid model experimentation is a priority
- Existing managed-service controls satisfy security requirements
- Dedicated GPU capacity would remain underutilised
Many enterprises will use a hybrid architecture.
Frequent and bounded interactions can run through a customer-hosted SLM. More complex cases can be routed to a larger approved model, a specialist workflow or a human agent. Consequential actions can pass through deterministic policy checks and approval controls regardless of which model produced the recommendation.
Where Astra SLM fits

Astra SLM is NuPlay’s 9B-parameter model for real-time enterprise voice and chat. It is designed for domain adaptation and deployment inside a customer’s VPC. NuPlay’s current product information describes it as multilingual, self-hosted and built for high-volume execution.
Its role in the wider voice architecture is specific:
- Understand the conversation
- Work with the context required for the current interaction
- Produce structured outputs for approved workflows
- Select from permitted tools
- Operate within the policies and controls surrounding the model
- Generate responses quickly enough for a natural spoken interaction
The deployment still includes the complete voice and execution stack around Astra. Telephony, speech processing, context retrieval, policy enforcement, tools, observability and human escalation all remain part of the architecture.
The value of Astra’s smaller footprint is that customer-VPC deployment becomes more practical. The value of the complete system depends on whether it meets the customer’s measured requirements for quality, latency, throughput, privacy, reliability and cost.
The bottom line
Self-hosting is an architectural choice, not a privacy guarantee.
An SLM is useful because its smaller computational footprint can make customer-controlled inference achievable for workloads that do not require a much larger model. In voice AI, that can support tighter data boundaries, greater control over dependencies and performance engineering, and more predictable economics at sustained scale.
Those benefits must be evaluated across the complete voice system. Model placement alone does not determine privacy, latency, security or reliability.
The right architecture is the one that meets the workload’s quality and operational requirements while giving the enterprise the degree of control it actually needs.
.gif)




.avif)
.avif)
.avif)
