08/10/2026

08/10/2026

08/10/2026

NeuDecide: An audio-first tiny decision model.

NeuDecide: An audio-first tiny decision model.

NeuDecide: An audio-first tiny decision model.

Choose a device. Tell it what to do - Interact with our demo below.

Choose a device. Tell it what to do - Interact with our demo below.

Select a device below, then record your voice. NeuDecide turns your words into actions.

Select a device below, then record your voice. NeuDecide turns your words into actions.

Select a device below, then record your voice. NeuDecide turns your words into actions.

Tell your devicewhat to do.RecordTRY SAYING“Vacuum the kitchen.”“Mop the bathroom.”“Return to yourdock.”“How much battery isleft?”“How full is thedustbin?”ModelTool selection{ }OutputLivingKitchenBedroomHallBathOfficeDOCK72%30%Robot vacuumiLEFTRIGHTRobot armi05:00SmartwatchiLampi
Robot vacuumRobot armSmartwatchLampTell your device what to do.RecordModelTool selection{ }STATUSPick a device, then press Record.The tool call the model returns shows up here.LivingKitchenBedroomHallBathOfficeDOCK72%30%Robot vacuumiTRY SAYING“Vacuum the kitchen.”“Mop the bathroom.”“Return to your dock.”“How much battery is left?”“How full is the dustbin?”

NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.

NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.

NeuDecide demo: Choose a device and tell it what to do. Your voice becomes action - no transcription needed.

We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.

Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.


NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.

We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.

Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.


NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.

We’re releasing a tiny audio-first decision model: at only 43mb size, running on a single thread of CPU, you can now embed voice interactions into all low powered hardware - including your browser.

Today’s approach to voice control is to transcribe the audio, pass the text to a tool-calling model, and validate the result before executing it. We wanted to see how much of that pipeline we could replace with a single small model.


NeuDecide takes audio and tool definitions as inputs and predicts a tool call directly, without the need of a transcription.

Decision models: AI that answers software, not people

Decision models: AI that answers software, not people

A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.

Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.


A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.

Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.


A decision model turns messy input into a bounded output that code can act on: a choice, a score, a function call. It never writes paragraphs. We decided to build this model after working closely with robotic partners who needed voice control that could run locally on limited hardware.

Since the TypeSafe release of their decision model Jev, we’ve seen many new releases and different interpretations of what a decision model should be. We’re exploring the same idea with speech as the input.


Why it has to listen

Why it has to listen

Why it has to listen

Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."

Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.


NeuDecide takes the audio directly.

Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."

Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.


NeuDecide takes the audio directly.

Real input is rarely clean text. An operator shouts across a warehouse. Someone mumbles to a smart phone on a windy street. A caller says "cancel it, no wait, just pause it."

Text-only decision models need speech recognition in front of them. That adds a second model, adds latency, and discards hesitation and emphasis.


NeuDecide takes the audio directly.

Built to listen

Built to listen

NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.


Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.

Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.

Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.

NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.


Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.

Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.

Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.

NeuDecide is an end-to-end voice-action model: a pre-trained speech encoder joined to a compact tool-calling encoder–decoder, then trained as one system.


Tool definitions. Tools are declared as JSON schemas, and the model answers with calls. Schemas are serialised exactly as Python's json.dumps writes them, and the vocabulary has dedicated <tools> and <tool_call> tokens. Training also used an auxiliary contrastive head alongside the decoder.

Design approach. Like Jev, NeuDecide treats a decision as something bounded, typed, and cheap enough to make thousands of times an hour. Jev's three primitives map neatly onto tool calls: a choice is which tool, a score is an argument, and null is the empty call.

Audio encoder. A streaming English speech encoder turns audio straight into the representations the decision model reads, so there is no transcript step in between.

How it works

How it works

NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.


NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.


NeuDecide is three ONNX graphs, 42.7 MB in total and roughly 60 million parameters. A request runs through them in order: encode the audio once, encode the tool list once, then decode the call one token at a time.



Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.


Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.


Only the small decoder runs in the loop, and it reuses cached keys and values, so each extra output token is cheap. Because the tool list is an input rather than baked into the weights, one model can operate a robot arm in the morning and a booking system in the afternoon. Changing the available actions therefore does not require any retraining. You change the JSON, not the model.

Built small from the start

Built small from the start

Built small from the start

NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.


The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.

NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.


The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.

NeuDecide was designed to be small from the start, rather than shrunk after training. Because the model learned to work at low precision during training, it keeps its accuracy at a fraction of the usual size.


The q4 export is 43 MB, and in our device benchmarks peak RAM stayed between 146 MB and 174 MB.

How it compares

How it compares

The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.


With a fraction of the parameters, NeuDecide gets more commands exactly right.


The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.


With a fraction of the parameters, NeuDecide gets more commands exactly right.


The standard way to build a voice agent today is a cascade: transcribe with a speech recogniser, then hand the text to a tool-calling model. We tested NeuDecide against cascades built from Parakeet, one of the most practical open speech recognition models, paired with the Needle and FunctionGemma tool-calling models.


With a fraction of the parameters, NeuDecide gets more commands exactly right.


Highlights

Highlights

Highlights

  • Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.

  • 97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.

  • 85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.

  • Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.

  • 94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.


Full results for every dataset, tool count and checkpoint are in the attached technical report.

  • Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.

  • 97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.

  • 85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.

  • Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.

  • 94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.


Full results for every dataset, tool count and checkpoint are in the attached technical report.

  • Model size: approximately 55M parameters: NeuDecide has around 2.5× fewer parameters than the 137M Parakeet + Needle cascade and 12.5× fewer than the 687M version.

  • 97.0% tool selection accuracy on Fluent Speech Commands: With ten tools available, NeuDecide selects the correct tool 97.0% of the time, compared with 90.1% for Parakeet 660M + Needle.

  • 85.3% argument accuracy on Fluent Speech Commands: With ten tools available, this compares with 69.4% for Parakeet 660M + Needle.

  • Similar accuracy with five and ten tools: On Neuphonic’s tool-only task, exact match changes from 95.0% to 94.8% as the tool count increases. The comparison cascade drops from 86.0% to 79.8%.

  • 94.8% exact match versus 47% for Parakeet + FunctionGemma: Both results are for Neuphonic’s tool-only task with ten tools.


Full results for every dataset, tool count and checkpoint are in the attached technical report.

What it costs to run

What it costs to run

We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.

We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.


We tested NeuDecide on a MacBook Pro M3, Samsung S24+ and Raspberry Pi 5, running on a single CPU thread. Time to call ranged from 46 ms to 206 ms, and loading took less than half a second on all three. Peak RAM ranged from 146 MB to 174 MB.

Results

Results

Results

NeuDecide performance on a single CPU thread

MacBook Pro M3 • Samsung S24+ • Raspberry Pi 5

Device

Device

Device

Time to call (ms)

Time to

call (ms)

Time to call (ms)

RTF

RTF

RTF

Peak RAM (MB)

Peak RAM
(MB)

Peak RAM (MB)

Loading time (ms)

Loading

time (ms)

Loading time (ms)

MacBook Pro M3

MacBook Pro M3

MacBook Pro M3

46

46

46

0.013

0.013

0.013

149

149

149

159

159

159

Samsung S24+

Samsung S24+

Samsung S24+

82

82

82

0.023

0.023

0.023

174

174

174

267

267

267

Raspberry Pi 5

Raspberry Pi 5

Raspberry Pi 5

206

206

206

0.058

0.058

0.058

146

146

146

499

499

499

What the numbers say

What the numbers say

What the numbers say

  • Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.

  • Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.

  • Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.

  • Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.

  • Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.

  • Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.

  • Time to call stayed below 210 ms across all three devices: 46 ms on MacBook Pro M3, 82 ms on Samsung S24+ and 206 ms on Raspberry Pi 5.

  • Loading took less than half a second: 159 ms, 267 ms and 499 ms respectively.

  • Peak RAM stayed below 175 MB: ranging from 146 MB to 174 MB across the tested devices.

Where it fits

Where it fits

NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.


NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.


NeuDecide is designed for products with a defined set of actions that users can trigger by voice, without sending audio to the cloud. Here are a few applications.


Call centres

Call centres

Call centres

This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.

  • Routing on the first sentence. Run the caller's opening words against tools like route_to(queue) and escalate(reason) before any transcript exists.

  • Live agent assist. Fire suggest_article, open_refund or verify_identity into the agent's screen while the customer is still talking.

  • Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.

  • Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.


This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.

  • Routing on the first sentence. Run the caller's opening words against tools like route_to(queue) and escalate(reason) before any transcript exists.

  • Live agent assist. Fire suggest_article, open_refund or verify_identity into the agent's screen while the customer is still talking.

  • Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.

  • Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.


This is the strongest fit today. A contact centre makes the same few hundred decisions millions of times a day: route to billing or retention, flag a complaint, open a refund, escalate to a person. And it runs on servers, where NeuDecide's memory use is no problem.

  • Routing on the first sentence. Run the caller's opening words against tools like route_to(queue) and escalate(reason) before any transcript exists.

  • Live agent assist. Fire suggest_article, open_refund or verify_identity into the agent's screen while the customer is still talking.

  • Every call, not a sample. At roughly a tenth of a second of one CPU thread per decision, it is cheap enough to run on all traffic.

  • Data stays home. Recordings never leave the operator's infrastructure, which simplifies GDPR and PCI reviews.


On-prem and regulated sites

On-prem and regulated sites

On-prem and regulated sites

Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.


Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.


Hospitals, banks, factories and government sites often cannot send voice data to an outside API. NeuDecide runs on an ordinary CPU through ONNX Runtime, with no GPU and no internet. A nurse says "log vitals for bed twelve" and the decision is made inside the building.


Robotics

Robotics

Robotics

A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.


A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.


A robot's action space is already a tool list: move_to(zone), pick(item), stop(), return_to_dock(). NeuDecide turns "grab the blue bin from aisle four" into pick with typed arguments, on the robot's own computer. Commands keep working when the Wi-Fi drops, and the planner can validate the call before a motor moves. On a Raspberry Pi 5, we measured 206 ms time to call and 146 MB peak RAM, running on a single CPU thread.


Wearables, next

Wearables, next

Wearables, next

Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.

In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.

Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.

In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.

Rings, earbuds and glasses are where audio-to-action matters most: "remind me when I get home" works offline, and raw audio never leaves the body. The files already fit. The runtime memory does not yet, so this use case waits on an int4-native runtime.

In every case, NeuDecide works best beside a larger model, not instead of one. It handles the fast, frequent, bounded decisions; anything it cannot place goes to a bigger model or a person.

Known limits

Known limits

A small, specialised model has sharp edges. These are the ones we found.

  • It can be trigger happy at times: Given 30 s of low-level noise, it still produced a set_timer call. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.

  • Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.

  • Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.

  • English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.

  • It only knows the tools you give it. Clear tool names and descriptions do most of the work.

  • Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.

A small, specialised model has sharp edges. These are the ones we found.

  • It can be trigger happy at times: Given 30 s of low-level noise, it still produced a set_timer call. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.

  • Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.

  • Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.

  • English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.

  • It only knows the tools you give it. Clear tool names and descriptions do most of the work.

  • Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.

A small, specialised model has sharp edges. These are the ones we found.

  • It can be trigger happy at times: Given 30 s of low-level noise, it still produced a set_timer call. Put a gate in front: voice-activity detection, a confidence threshold, or an explicit "no action" tool. This export has no calibrated confidence score to threshold on, so the gate has to sit outside the model.

  • Memory depends on the runtime and workload. Model files total 43 MB; peak RAM ranged from 146 MB to 174 MB in our device tests.

  • Long tool lists become inaccurate: We recommend a maximum of 10 tools for accuracy.

  • English only, 30 seconds max. Inputs are capped at 30 s of audio, 1,536 tool tokens and 128 output tokens.

  • It only knows the tools you give it. Clear tool names and descriptions do most of the work.

  • Open-ended names are harder. On SNIPS, where arguments are often free-form names such as artists and playlists, cascades that work from a transcript still fill arguments more accurately. Test it on your own tools and your users' voices before deploying.

What comes next

What comes next

For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.

Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.

For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.

Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.

For many tasks, AI inside a product does not need to chat. NeuDecide shows that it does not need a transcript either: it can turn speech directly into a tool call. In our evaluations, one small model that listens and decides beats a chain of large ones.

Next, we’re working on streaming inference, grammar-constrained decoding, and a confidence head to help applications decide when to escalate. We’re also exploring how to bring NeuDecide to smaller devices with tighter memory and power budgets. We’re also improving argument accuracy on harder datasets such as SNIPS.

Sources

Sources

  • TypeSafe AI: Jev, ThursdAI release coverage

  • NeuDecide config.json and ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)

  • TypeSafe AI: Jev, ThursdAI release coverage

  • NeuDecide config.json and ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)

  • TypeSafe AI: Jev, ThursdAI release coverage

  • NeuDecide config.json and ONNX graphs (audio encoder, tool encoder, decoder step, q4f variant)

Relevant Links

Relevant Links

Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide

Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide

Get the Python package: https://github.com/neuphonic/neudecide

Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide

Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide

Get the Python package: https://github.com/neuphonic/neudecide

Download the model from Hugging Face: https://huggingface.co/neuphonic/neudecide

Try out the demo on the Hugging Face space: https://huggingface.co/spaces/neuphonic/neudecide

Get the Python package: https://github.com/neuphonic/neudecide

#DecisionModels

#DecisionModels

#DecisionModels

#AudioToAction

#AudioToAction

#AudioToAction

#ToolCalling

#ToolCalling

#ToolCalling

Bring Neuphonic into your product.

Building something with voice AI? Talk to us about your use case,

deployment requirements, and how Neuphonic can fit into your stack.

Building something with voice AI? Talk to us about your use case,

deployment requirements, and how Neuphonic can fit into your stack.