08/10/2026
08/10/2026
08/10/2026
NeuDecide: An audio-first tiny decision model.
NeuDecide: An audio-first tiny decision model.
NeuDecide: An audio-first tiny decision model.
Choose a device. Tell it what to do - Interact with our demo below.
Choose a device. Tell it what to do - Interact with our demo below.
Select a device below, then record your voice. NeuDecide turns your words into actions.
Select a device below, then record your voice. NeuDecide turns your words into actions.
Select a device below, then record your voice. NeuDecide turns your words into actions.
NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.
NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.
NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.
We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.
Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.
NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.
We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.
Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.
NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.
We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.
Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.
NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.
Decision models: AI that answers software, not people
Decision models: AI that answers software, not people
A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.
Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.
A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.
Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.
A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.
Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.
Why it has to listen
Why it has to listen
Why it has to listen
Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."
Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.
NeuDecide takes the audio directly.
Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."
Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.
NeuDecide takes the audio directly.
Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."
Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.
NeuDecide takes the audio directly.
Built to listen
Built to listen
NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.
Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.
Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.
Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.
NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.
Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.
Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.
Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.
NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.
Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.
Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.
Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.
How it works
How it works
NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.
NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.
NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.

Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.
Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.
Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.
Built small from the start
Built small from the start
Built small from the start
NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.
The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.
NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.
The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.
NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.
The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.
How it compares
How it compares
The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.
With a fraction of the parameters, NeuDecide gets more commands exactly right.
The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.
With a fraction of the parameters, NeuDecide gets more commands exactly right.
The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.
With a fraction of the parameters, NeuDecide gets more commands exactly right.

Highlights
Highlights
Highlights
Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.
97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.
85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.
Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.
94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.
Full results for every dataset, tool count and checkpoint are in the attached technical report.
Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.
97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.
85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.
Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.
94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.
Full results for every dataset, tool count and checkpoint are in the attached technical report.
Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.
97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.
85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.
Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.
94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.
Full results for every dataset, tool count and checkpoint are in the attached technical report.
What it costs to run
What it costs to run
We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.
We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.
We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.
Results
Results
Results
NeuDecide performance on a single CPU thread
MacBook Pro M3 • Samsung S24+ • Raspberry Pi 5
Device
Device
Device
Time to call (ms)
Time to
call (ms)
Time to call (ms)
RTF
RTF
RTF
Peak RAM (MB)
Peak RAM
(MB)
Peak RAM (MB)
Loading time (ms)
Loading
time (ms)
Loading time (ms)
MacBook Pro M3
MacBook Pro M3
MacBook Pro M3
46
46
46
0.013
0.013
0.013
149
149
149
159
159
159
Samsung S24+
Samsung S24+
Samsung S24+
82
82
82
0.023
0.023
0.023
174
174
174
267
267
267
Raspberry Pi 5
Raspberry Pi 5
Raspberry Pi 5
206
206
206
0.058
0.058
0.058
146
146
146
499
499
499
What the numbers say
What the numbers say
What the numbers say
Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.
Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.
Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.
Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.
Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.
Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.
Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.
Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.
Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.
Where it fits
Where it fits
NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.
NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.
NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.
Call centres
Call centres
Call centres
This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.
Routing on the first sentence. Run the caller's opening words against tools like
route_to(queue)andescalate(reason)before any transcript exists.Live agent assist. Fire
suggest_article,open_refundorverify_identityinto the agent's screen while the customer is still talking.Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.
Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.
This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.
Routing on the first sentence. Run the caller's opening words against tools like
route_to(queue)andescalate(reason)before any transcript exists.Live agent assist. Fire
suggest_article,open_refundorverify_identityinto the agent's screen while the customer is still talking.Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.
Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.
This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.
Routing on the first sentence. Run the caller's opening words against tools like
route_to(queue)andescalate(reason)before any transcript exists.Live agent assist. Fire
suggest_article,open_refundorverify_identityinto the agent's screen while the customer is still talking.Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.
Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.
On-prem and regulated sites
On-prem and regulated sites
On-prem and regulated sites
Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.
Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.
Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.
Robotics
Robotics
Robotics
A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.
A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.
A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.
Wearables, next
Wearables, next
Wearables, next
Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.
In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.
Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.
In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.
Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.
In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.
Known limits
Known limits
A small, specialised model has sharp edges. These are the ones we found.
It can be trigger happy at times: Given 30 s of low-level noise, it still produced a
set_timercall. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.
Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.
English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.
It only knows the tools you give it. Clear tool names and descriptions do most of the work.
Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.
A small, specialised model has sharp edges. These are the ones we found.
It can be trigger happy at times: Given 30 s of low-level noise, it still produced a
set_timercall. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.
Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.
English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.
It only knows the tools you give it. Clear tool names and descriptions do most of the work.
Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.
A small, specialised model has sharp edges. These are the ones we found.
It can be trigger happy at times: Given 30 s of low-level noise, it still produced a
set_timercall. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.
Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.
English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.
It only knows the tools you give it. Clear tool names and descriptions do most of the work.
Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.
What comes next
What comes next
For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.
Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.
For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.
Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.
For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.
Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.
Sources
Sources
Needle 2 model card, Cactus Compute on Hugging Face
TypeSafe AI: Jev, ThursdAI release coverage
NeuDecide
config.jsonand ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)
Needle 2 model card, Cactus Compute on Hugging Face
TypeSafe AI: Jev, ThursdAI release coverage
NeuDecide
config.jsonand ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)
Needle 2 model card, Cactus Compute on Hugging Face
TypeSafe AI: Jev, ThursdAI release coverage
NeuDecide
config.jsonand ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)
Relevant Links
Relevant Links
Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide
Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide
Get the Python package: https://github.com/neuphonic/neudecide
Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide
Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide
Get the Python package: https://github.com/neuphonic/neudecide
Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide
Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide
Get the Python package: https://github.com/neuphonic/neudecide
#DecisionModels
#DecisionModels
#DecisionModels
#AudioToAction
#AudioToAction
#AudioToAction
#ToolCalling
#ToolCalling
#ToolCalling
Bring Neuphonic into your product.
Building something with voice AI? Talk to us about your use case,
deployment requirements, and how Neuphonic can fit into your stack.
Building something with voice AI? Talk to us about your use case,
deployment requirements, and how Neuphonic can fit into your stack.
Neuphonic © 2026

Neuphonic © 2026

Neuphonic © 2026

