Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

92% Positive

Analyzed from 980 words in the discussion.

Trending Topics

#read#preadv#mmap#memory#cache#fast#faster#performance#kernel#page

Discussion (35 Comments)Read Original on HackerNews

ComputerGuru•about 14 hours ago
Interesting article but it gave me a bit of a panic attack. Benchmarking (with TPC or otherwise) is NOT the way to determine the correct approach here; that is strictly only to be used for databases (typically RDBMS) effectively “owning” the complete hardware they are running on. An embedded database might be used in that manner if it’s operating as the backend for a pure crud application that performs ~zero server side rendering, parsing, validation, etc and is essentially just an async http-to-SQLite interface. But more likely than not, an embedded db will be used and deployed on machines (not necessarily even servers) serving many a purpose, and need to perform best both within the confines of the resources available to the machine and in relative terms, necessarily making tradeoffs that might sacrifice performance for “value” in terms of CPU or memory usage.

This isn’t just with regards to benchmarking, it’s an essential consideration *any* time you are taking ownership of the cache away from the kernel, which is the only piece in the stack that has viability into the global state and can be trusted to give back memory under pressure to ensure everything plays nice together. It’s not limited to just databases or even just memory, for example FreeBSD has had greater than its fair share of issues that trace back to the ZFS having a separate cache from the kernel (despite the much tighter integration between the two and presence of various mechanisms to address pathological cases). You can also refer to any comparison or benchmark between the use of spinlocks vs mutexes: spin locks consistently perform better in/on (micro)benchmarks but are almost always actually the worse choice in the grand scheme of things because the benchmarks falsely assume complete and uncontended ownership over system resources.

This isn’t even io_uring specific and I’m hardly the first to bring this up in the context of O_DIRECT.

shellpipe•about 10 hours ago
TPC-H is the recommended approach in Turso's CONTRIBUTING.md - https://github.com/tursodatabase/turso/blob/main/CONTRIBUTIN...
vlovich123•about 6 hours ago
You’ve misunderstood what they’re saying. Not that the benchmark is invalid, but that it’s an incomplete picture if your use case is a mixed set of applications where DB performance is not the only important thing. Hence the comment about memory - O_DIRECT is optimal in terms of DB performance but then the caches are owned by the application and the kernel isn’t free to discard those caches whenever it wants like it can with the page cache. I’m not actually sold on this argument though - I have an embedded DB that with a very tiny block cache goes very very fast, and uses far less RAM than using the page cache because it knows when read ahead is called for vs when it’s not
pbowyer•about 11 hours ago
What would be the correct way to do benchmarking/profiling here? I ask as I've started to run some for Vinyl cache to see what performance issues I can find, and I've quickly learned just how hard it is to get good, repeatable benchmarks that isolate the right thing.
marginalia_nu•about 16 hours ago
If you're doing contiguous readahead in userspace, why not just use preadv? It'll limit you to doing readahead up until the next resident page, but at least in my experiments in Marginalia's index, preadv beats io_uring in all cases you can use a single preadv call to do the full read.
vlovich123•about 16 hours ago
Not sure what you mean. Nothing about preadv lets you indicate you only want to read what's already in the page cache. And io_uring and preadv aren't orthogonal - you can give io_uring a preadv op to do the scattered read instead of issuing separate read OPs although I'm not 100% sure how much of a win that is in practice.

Also, I think you misunderstood the blog as it's describing application read-ahead which is what you have to do when using O_DIRECT.

wahern•about 13 hours ago
> Nothing about preadv lets you indicate you only want to read what's already in the page cache.

preadv2 + RWF_NOWAIT

marginalia_nu•about 16 hours ago
So I have a buffer pool with O_DIRECT reads.

I implement read-ahead in the application by (optionally) preadv:ing a single read into multiple destination buffers in the pool, leaving them unpinned, since as long as you aren't up against the bandwidth limit of the drive, a larger read is generally as fast as multiple smaller one on modern hardware.

I've tried doing this with io_uring as well, but found just eating the preadv syscall cost was faster.

vlovich123•about 13 hours ago
Do this with io_uring with the preadv syscall. It’ll be the same or faster (faster only if you can do something else while waiting for I/O or you can submit multiple requests simultaneously - a single io_uring will be basically identical)
lossolo•about 15 hours ago
> a larger read is generally as fast as multiple smaller one on modern hardware.

Not always if by modern you mean NVMe drives. One synchronous preadv() for 256 KiB gives the kernel/device one big request but 16 independent asynchronous 16 KiB reads can be serviced concurrently. So the latter gives the NVMe controller 16 operations it can schedule in parallel. So depending on the workload and hardware, offsets, filesystem and request sizes that can give you lower aggregate latency or higher throughput.

Asmod4n•about 15 hours ago
I found out that using mmap and just telling uring to read form there to beat anything else.
marginalia_nu•about 15 hours ago
Depends a lot on the memory pressure. If you can be fairly certain the data is (or will be) resident in memory, mmap is basically unbeatable. If you can't (because the data is larger than RAM or there's other stuff competing for RAM), mmap can have gnarly system-wide performance implications[1].

[1] Mandatory mmap=poop-emoji link: https://db.cs.cmu.edu/mmap-cidr2022/

Asmod4n•about 14 hours ago
In my case: its a PROT_READ, MAP_PRIVAT map and i cannot get a SIGBUS since i tell the kernel to handle it all for me, thanks to liburing, instead i get a short send in that case.
megagpt1•about 15 hours ago
Databases vendors usually find that read outperforms mmap. Mmap being fast is a myth.
Asmod4n•about 12 hours ago
Tried every way to get contents of a file as fast to a client socket as possible.

mmap and io_uring_prep_send were faster than everything else, no matter the size as long as you keep the map around for the lifetime of the process.

for one off sends when a file is smaller than 256kb then io_uring_prep_read + prep_send are faster than everything else.

yxhuvud•about 13 hours ago
Note that grandparent actually didn't say they used mmap to actually get the memory, just mapping it.
0c3ca83•about 15 hours ago
Why would you use uring to read from an mmap? Couldn't you just memcpy?
Asmod4n•about 14 hours ago
To be more precise: i use that map to send assets out directly to clients from a zip file.

Its a new web server i am building and its the fastest way i could find out.

Just switching from epoll to liburing made the server ~45% faster too, its ridiculous. It can serve 10 gigabyte per second with a single thread, or around 10 million responses per second with h2 and 32 multiplexed requests.

I had to write a new http load generator for that since i couldn't find one which could generate enough load to saturate my server or be fast enough to withstand it.

jeffbee•about 15 hours ago
Maybe they meant using IORING_OP_MADVISE with MADV_WILLNEED to bring ranges into memory?
amluto•about 15 hours ago
I’m curious why the choice is between syscalls and, specifically, io_uring with O_DIRECT. AFAIK Turso is like SQLite and supports multiple processes accessing the same database, and I would expect buffering to be a huge win in some workloads. What’s wrong with io_uring without direct? There’s also the middle ground of RWF_DONTCACHE.
ErroneousBosh•about 11 hours ago
This has scrolled past in my feed and every time I've read it as "Io_uring without Radiohead", and I mean you could but what would be the point?