Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This reminds me of a horrible hack I once put into an internal ruby (sinatra) REST API server that had to aggregate data from a ton of SOAP servers in a pretty large environment, and return some computed results as JSON. The amount of XML that had to be parsed into objects to do the reporting was massive, and the web workers (being ruby) never gave the memory back to the OS. Running the code inside a unicorn process would basically cause it to start using several GB of ram, and over time all the unicorn workers would become bloated and the service would have to be restarted.

I eventually decided to just fork() the ruby process, do the processing, compute the JSON results (which were comparatively small) and hand the results back as a string over a socket to the parent process, which would cache them and respond. It pretty much solved the memory issue, since the forked process would just exit() gave all the memory back.

In fact, I think I ended up doing a GC.disable() inside the forked process (because I was going to exit anyway) and it sped the process up from taking >2 minutes to just a few seconds. Ruby's GC was utterly terrible back then.

It was an internal tool that never was touched by the general internet, which is why we could "afford" to have it do such intensive work behind a single request, but I always wonder if there was a better approach that wasn't so awful.



Ah, the "missile" approach to memory management: https://groups.google.com/forum/message/raw?msg=comp.lang.ad...


Fun fact: Erlang is built to assume that most processes are "missiles" in this sense. GC is quite slow (though only blocks the individual actor being collected), but most actors never perform enough allocations+deallocations across their lifetime to ever incur a GC collection.

One of the main ways to tune Erlang systems for speed is to figure out the amount of memory each type of actor will need, and pass an option to spawn the actor with an arena pre-allocated to that size, so that it can just quietly do its job and quit, never having asked to allocate or deallocate at all.


hmm www.rational.com is broken, but is that the Rational from Rational Unified Process?


Yes. And, may god have mercy on its soul, Rational ClearCase. One of the more unhelpful proprietary version control systems.

https://en.wikipedia.org/wiki/Rational_Software


It's fairly common to have workers limited to a number of requests or memory size for similar reasons, forking new ones regularly.

Instagram found deactivating parts of Python's memory management actually helped with efficiency in such a scenario: https://instagram-engineering.com/dismissing-python-garbage-...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: