UltrafastSecp256k1: A High-Performance secp256k1 Execution Engine for Wallets, Privacy, and Scalable BCH Applications

@shrec re: privacy and reusable addresses, the focus has been on how to do that in a post quantum (pq) resistant fashion, which is currently quite challenging.

since rpa is based on ecdh, it’s not currently seen a viable implementation for privacy in pq era and the question is now what to use instead? the challenge there is also, which offer reusable addressing?

some food for thought

That’s fair. If the requirement is “post-quantum reusable addressing as the final design”, then ECDH-based RPA is not the endgame.

But I think these are two different timelines.

Today, BCH still uses secp256k1. If secp256k1 becomes practically unsafe, the problem is much larger than RPA: signatures, wallets, existing coins, address reuse assumptions, and migration rules all need to change.

Until that migration path exists, I think current-era privacy/reusable-address designs are still worth evaluating. RPA may not be the final PQ answer, but it can still be a practical workload for today’s BCH ecosystem — especially for mobile/L1 usability.

So I’d separate the questions:

  1. What can improve privacy and reusable-address UX today under the current secp256k1 model?
  2. What should replace that model in a future PQ migration?

Both are important, but I wouldn’t block current privacy work on a PQ design that the ecosystem cannot yet deploy or test in practice.

Thanks for tagging Fernando.

For Knuth specifically, I already built a shim and ran an A/B benchmark using Knuth’s own tooling. My goal is not to ask for a default replacement, but to provide a clean optional evaluation path: build it, run the same benchmarks, inspect the shim, and decide based on reproducible results.

I’ll prepare a Knuth-specific integration note after the v3.69 release so the benchmark can be reproduced independently.

yes, it’s a complex question. there has been extensive work evaluating this from an RPA perspective https://github.com/00-Protocol/BCH-Stealth-Protocol/tree/main

and in my thread above, there is a pq path towards privacy demonstrated on chipnet today where amounts are hidden. if paired with cashfusion, identity may also be shielded.

so it’s not so much that pq is blocking privacy development as it is how are the questions being created by that work getting answered? for example, reusable addresses with ml-kem-768

I personally didn’t run into any performance issues with secp256k1. what have you experienced?

regarding the RPA “grind” I believe the bottleneck there is electrum/fulcrum performance

Thanks, that makes sense — I’ll read through the BCH Stealth Protocol work.

I agree that PQ privacy/reusable-address research is important, especially if the goal is a long-term design rather than only a current-era secp256k1 design.

The distinction I’m trying to make is mostly about timelines and deployment targets:

  1. Current production BCH still runs on secp256k1, so RPA/ECDH-style workloads can still be useful to evaluate for today’s wallet/mobile/L1 performance and privacy UX.

  2. PQ reusable addressing is a separate research/migration track. ML-KEM-768 or hybrid designs may be the right direction, but that needs its own engineering work: address format, key size/UX, wallet scanning model, compatibility, migration, and testnet/chipnet validation.

So I don’t see these as mutually exclusive. RPA may not be the final PQ answer, but it can still be a useful current-era workload. In parallel, PQ reusable addressing should be explored as the future-facing track.

For my library, the immediate relevance is: if a design has heavy EC workloads today, I can help benchmark/optimize that. If the future design uses ML-KEM or hybrid primitives, that’s probably a separate module/track rather than something I would mix into the secp256k1 engine directly.

Fair point. I’m not claiming that secp256k1 is currently the main BCH bottleneck in normal node operation.

What I have so far are A/B integration benchmarks showing that the same secp256k1 workloads can be executed significantly faster through the shim in several existing codebases. That does not automatically mean those workloads are the current production bottleneck.

For RPA specifically, if the main bottleneck is Electrum/Fulcrum indexing/scanning rather than raw EC operations, that is useful information. Then the right next step is not to claim “secp256k1 is the bottleneck”, but to build a workload-level benchmark:

  • wallet-side RPA scan/grind cost
  • Electrum/Fulcrum query/indexing cost
  • EC operation cost inside the pipeline
  • mobile/client-side constraints
  • batching opportunities

My interest is mostly in identifying where faster EC execution becomes useful, not forcing it into places where it is not the limiting factor.

So if you have a good RPA benchmark path or a pointer to where Fulcrum/Electrum spends time, I’d be interested in testing that directly.

2 Likes

A 60-80% signing time reduction is quite significant! The problem with novel cryptographic libs in general, though, is safety: The engine is so central to everything, people would likely want not just intense scrutiny but also a proven track record… The track record part faces a harsh chicken and egg problem.

As others kind of mentioned up there, RPA may help you solve the chicken and egg problem of track records, since it’s a sufficiently small niche yet is significant enough to attract some significant attention. The grind may not the current biggest bottleneck, but then the current implementation is nowhere near ideal and neither is the grind target bits - imo it should be grinding more bits given current traffic so the indexer doesn’t need to transmit as much stuff.

This is probably a huge ask, but if you feel confident enough, the best path forward might be to directly help out the RPA project even for the parts not directly related to your lib. More popularity, less unrelated bottlenecks, higher likelihood for more track record.

3 Likes

That makes sense. After I finish the current release/PR package, I’m planning to look at the full RPA pipeline directly.

I don’t want to treat it as only an EC primitive problem. I’d rather measure the whole path, identify the real bottlenecks, optimize what can be optimized, and then return a working, reproducible solution with benchmarks and documentation.

If the bottleneck is the indexer, that should be improved. If the EC layer matters, that should be measured too. The useful result is not a claim — it is a better RPA pipeline that others can test.

3 Likes

The way that most BCH end users might end up accessing secp256k1 is through a library called libauth, which makes a C implementation of secp256k1 available over wasm.

The secp256k1 implementation is here:

That’s the secp256k1 implementation underneath cashonize.com, mainnet.cash, selene.cash, CashScript.org, cauldron.quest, anyhedge, vox.cash, etc.

There are some old benchmarks here:

The libauth repo is configured to check benchmarks in the CI, and there’s yarn commands set up to do it.

To my knowledge, nobody has implemented RPAs on a web/webkit/electron platform.

1 Like

i will finish release 4.0 and will focus on rpa will find solution. my engine on bitcoin made silent payment possible frigate electrum works on my engine 3.68 version

1 Like

There is a unique application for this on Bitcoin Cash now.

We have a minable CashToken (& price oracle) called Photons which uses an secp256k1 signing as part of the mining algorithm. With very rudimentary in browser profiling, it appears the Schnorr signing consumes roughly 80% of the tread time with the wasm libauth implementation and a CPU.

As your library appears to be the fastest on the market, ChatGPT is directing miners to your implementation as a more efficient alternative. The GPU implementation is roughly a factor of 70 times more efficient than our old libauth implementation.

Is there any way to make this library work with WebGPU or WebGL?

if possible i will found solution just i need test from your side to make me test repo where i will integrate this and test so if you make me standart test repo that i will fork and change later to make tests i will try my best make me small repo and documantation what part of operations i will replace and all other i will tell after that

what i see i think it’s possible with WASM + WebGPU i need your repo that i will acelarate wasm i have already but in wasm i have not added gpu support before but will be interesting to work on it from my side

For CUDA and CPU mining with your libraries, I added some of the benchmark files for mining photons to a repo here:

Some links in the c++ may be broken as I moved them from a top-level in another copy of the repo.

The above was made by ChatGPT pro, I have not run or verified it myself. I would NOT trust the BCH Secp256 implementation anywhere outside the current application, where it seems to work quite well.

Your issue seems to give a straight-forward proposal for exploring the topic.

@shrec Thanks for publishing this work it helped me in building a high-performance optimized GPU miner for Cashtokens

2 Likes

you using my engin on backend or based my codes you build your code separatly ? if my engin is used i will add your project in adoptions

Thanks, @ABLA. I’m glad UltrafastSecp256k1 helped you build Pickaxe, and I appreciate the work you’ve put into the miner.

I’ve been optimizing a local version and am now seeing about 1.22 GH/s on my RTX5060ti GPU. I’d prefer to work together on bringing those improvements to Pickaxe. If you’re interested, let’s discuss the integration and a fair share of Pickaxe’s existing donation privately.

I integrated a Rust port of UltrafastSecp256k1 CPU and NVIDIA GPU arithmetic into Pickaxe and contributed it back through PR #444.

Tests ran on an NVIDIA RTX 50-series GPU with roughly desktop RTX 5070 Ti- Laptop class graphics performance:

Implementation Hash rate Comparison
PHOTON reference associated with 2qx’s work ~2 MH/s, estimated Reference baseline
My legacy Pickaxe kernel 114.24 MH/s ~57× the estimated reference
Pickaxe with the UltrafastSecp256k1 Rust port 123.17 MH/s 7.82% above my legacy kernel

The initial direct integration of UltrafastSecp256k1 produced a lower hash rate for Pickaxe’s workload. I therefore ported the required arithmetic into Rust while retaining Pickaxe’s specialized mining logic and optimizations.

Repeated tests on the same GPU showed 6.6–7.8% higher hash rate, with correctness checks passing. I selected this backend for the updated v0.0.1 and kept my legacy implementation available for reference.

1 Like

Confirmed! The leap in performance is unprecedented, and this is truly amazing work!

4 Likes