cosmopolitan

mirror of https://github.com/jart/cosmopolitan.git synced 2025-07-02 09:18:31 +00:00

Author	SHA1	Message	Date
Justine Tunney	3609f65de3	Make malloc() go 200x faster If pthread_create() is linked into the binary, then the cosmo runtime will create an independent dlmalloc arena for each core. Whenever the malloc() function is used it will index `g_heaps[sched_getcpu() / 2]` to find the arena with the greatest hyperthread / numa locality. This may be configured via an environment variable. For example if you say `export COSMOPOLITAN_HEAP_COUNT=1` then you can restore the old ways. Your process may be configured to have anywhere between 1 - 128 heaps We need this revision because it makes multithreaded C++ applications faster. For example, an HTTP server I'm working on that makes extreme use of the STL went from 16k to 2000k requests per second, after this change was made. To understand why, try out the malloc_test benchmark which calls malloc() + realloc() in a loop across many threads, which sees a a 250x improvement in process clock time and 200x on wall time The tradeoff is this adds ~25ns of latency to individual malloc calls compared to MODE=tiny, once the cosmo runtime has transitioned into a fully multi-threaded state. If you don't need malloc() to be scalable then cosmo provides many options for you. For starters the heap count variable above can be set to put the process back in single heap mode plus you can go even faster still, if you include tinymalloc.inc like many of the programs in tool/build/.. are already doing since that'll shave tens of kb off your binary footprint too. Theres also MODE=tiny which is configured to use just 1 plain old dlmalloc arena by default Another tradeoff is we need more memory now (except in MODE=tiny), to track the provenance of memory allocation. This is so allocations can be freely shared across threads, and because OSes can reschedule code to different CPUs at any time.	2024-06-05 02:02:14 -07:00
Jōshin	6e6fc38935	Apply clang-format update to repo (#1154 ) Commit `bc6c183` introduced a bunch of discrepancies between what files look like in the repo and what clang-format says they should look like. However, there were already a few discrepancies prior to that. Most of these discrepancies seemed to be unintentional, but a few of them were load-bearing (e.g., a #include that violated header ordering needing something to have been #defined by a 'later' #include.) I opted to take what I hope is a relatively smooth-brained approach: I reverted the .clang-format change, ran clang-format on the whole repo, reapplied the .clang-format change, reran clang-format again, and then reverted the commit that contained the first run. Thus the full effect of this PR should only be to apply the changed formatting rules to the repo, and from skimming the results, this seems to be the case. My work can be checked by applying the short, manual commits, and then rerunning the command listed in the autogenerated commits (those whose messages I have prefixed auto:) and seeing if your results agree. It might be that the other diffs should be fixed at some point but I'm leaving that aside for now. fd '\.c(c\|pp)?$' --print0\| xargs -0 clang-format -i	2024-04-25 10:38:00 -07:00
Justine Tunney	2ab9e9f7fd	Make improvements - Introduce portable sched_getcpu() api - Support GCC's __target_clones__ feature - Make fma() go faster on x86 in default mode - Remove some asan checks from core libraries - WinMain() now ensures $HOME and $USER are defined	2024-02-12 10:23:00 -08:00
Justine Tunney	5f8e9f14c1	Add OpenMP support	2024-01-28 22:39:02 -08:00

4 commits