My journey optimizing a Django application

  • python
  • django
  • gunicorn
  • rinha de backend
My journey optimizing a Django application

Rinha de backend is a programming challenge aimed at generating and sharing knowledge. This post describes my experience participating in the challenge, but I believe it is useful for anyone who wants to understand a bit more about Django, gunicorn and gevent.

For the second edition of the rinha de backend I made a submission using Django with the goal of improving my knowledge of Python and the framework. The challenge consists of solving the classic bank transaction concurrency problem while supporting a load test. The code resulting from my submission can be seen here.

Besides Django, I chose to use:

  • Django REST Framework
  • Gunicorn as the web server
  • PostgreSQL as the database

My first goal was to run the API load tests without resource restrictions, using all the resources available on my computer. After that, to submit my solution to the challenge, it was necessary to apply CPU (1.5) and memory (550MB) limits. That means the sum of the resources of the load balancer, database and two API instances had to stay within the specified limit — and, if possible, keep request times low throughout the whole test.

It is worth mentioning that most of the optimizations were made specifically for the rinha and are not necessarily needed — or even recommended — in real applications. At the same time, I feel much more comfortable putting Django applications in production with the knowledge I acquired during this journey.

Gevent

Gevent is a Python coroutine library that makes it easier to run blocking code asynchronously. An example of the impact of using gevent can be seen in this article. Since I have more experience with C#, I draw a parallel with async/await even though they have distinct implementations.

Gunicorn offers different worker types and the documentation recommends:

Some examples of behavior requiring asynchronous workers: Applications making long blocking calls (Ie, external web services)

Since the API is basically IO bound, gevent was chosen as the worker type.

Micro-optimizations

Since one of the goals was to have an optimized API, I searched for and found the following topics to improve my code’s performance:

  • values(): returns a QuerySet that returns dictionaries instead of model instances
  • update_fields: an argument of the save method that specifies which fields should be updated, simplifying the generated SQL command
  • Removing unnecessary middlewares (for the challenge)

I did not measure the impact but kept the changes. Based on the documentation and blog posts, some performance gain is expected.

Some things I tried but abandoned:

  • Pgbouncer — it solved the Postgres connection limit problem. But the “healthy” API kept the number of connections low enough.
  • Pypy instead of CPython — the results were not good: it increased resource consumption while delivering worse performance. That said, it is important to note that I did not dig into the subject and spent no time on any tuning.
  • Using psycopg3 directly with its connection pool. I got great results this way, and it showed that the Django + gunicorn + gevent setup works well and is able to process the requests within the rinha’s resource limitations. But I still wanted a submission using the ORM.

Persistent connections (CONN_MAX_AGE)

Django does not have its own connection pool for database connections. One possibility is the CONN_MAX_AGE property, which determines the lifetime a connection can have. When 0 (the default value), each connection is closed at the end of the request execution. Values greater than 0 indicate how long a connection can exist before being closed, allowing other requests to reuse the connection. In practice I noticed that using any value other than 0 does not work well with gunicorn’s gevent worker type. As described here, connections are not reused in this scenario.

With CONN_MAX_AGE!=0 and worker_type=gevent there is no reuse of persistent connections. The total number of connections grows even with almost all connections idle. As the number of requests per second increases, the number of connections also grows With CONN_MAX_AGE!=0 and worker_type=gevent there is no reuse of persistent connections. The total number of connections grows even with almost all connections idle. As the number of requests per second increases, the number of connections also grows

Keeping this value at 0 keeps the number of Postgres sessions low, with no need to reuse connections for that purpose.

With CONN_MAX_AGE=0 and worker_type=gevent the number of connections remained stable. Each request opens and closes a new database connection With CONN_MAX_AGE=0 and worker_type=gevent the number of connections remained stable. Each request opens and closes a new database connection

Worker connections

The default value of worker connections is 1000, so starting the tests with the value 50 and increasing from there seemed reasonable. However, while using that “high” value, during the test the number of Postgres sessions started stable until reaching a peak moment and from that point on the application was no longer able to respond to the load test requests.

With a higher value of GUNICORN_WORKER_CONNECTIONS and GUNICORN_WORKER_TYPE=gevent, at a certain point during the load test the number of database connections exploded, causing API errors With a higher value of GUNICORN_WORKER_CONNECTIONS and GUNICORN_WORKER_TYPE=gevent, at a certain point during the load test the number of database connections exploded, causing API errors

I accidentally ran a test without gevent and, to my surprise, the tests completed without errors. Reading this article I understood the reason. The high number of worker connections combined with gevent meant that more greenlets (pseudo threads) were scheduled beyond the processing capacity. A request would start and, due to a blocking operation such as a database query, its execution would be paused and sent to the queue of items to be processed. Since several requests were already ahead, that greenlet suffered starvation, and this process causes a spike in the number of open sessions when a sufficient load is applied. Adjusting this value to a lower number (5 in my case) already shows a considerable improvement in the test results.

With worker_connections=5 the database connections remained low during the whole test With worker_connections=5 the database connections remained low during the whole test.

Part of the Gatling load test execution report. The result shows good performance but the resource limits were not applied yet Part of the Gatling load test execution report. The result shows good performance but the resource limits were not applied yet

Profiling

Although very happy with the API’s execution, I was still skeptical because the metrics collected during the tests showed that the API would not behave well under the challenge’s imposed limits, especially for CPU.

Result of the docker stats command. CPU values of the APIs exceed (by a lot) the challenge limits Result of the docker stats command. CPU values of the APIs exceed (by a lot) the challenge limits

When limited, the load test results showed a high error rate because the API could not process the requests in time to respond to the load being sent. To try to understand the reason, I used a very interesting profiling tool that won me over with its ease of use. With a simple command, py-spy let me generate a flamegraph of my API during the load and see which methods consume the most CPU time.

py-spy record --gil --subprocesses -o profile.svg -- gunicorn -w 2 rinha.wsgi -b 0:9999

Flamegraph of a worker generated by py-spy during the load test execution Flamegraph of a worker generated by py-spy during the load test execution

Flamegraph focused on the database transaction opening Flamegraph focused on the database transaction opening

The generated file is a navigable SVG and can be downloaded here.

Back to CONN_MAX_AGE

On the last day of the rinha, a few hours before the end, while looking at the graph I noticed that a lot of CPU time was spent creating the database connection. So I speculated that if there was a connection pool, or at least a way to avoid opening connections on every request, CPU consumption would be reduced. That brought back the idea of setting a CONN_MAX_AGE value again to keep connections open and reused. But for that it would be necessary to change the worker type to sync. The expectation was that the lower resource consumption would allow the API to run under load with the imposed limitations. And fortunately that is what happened.

Part of the Gatling load test execution report. The result shows good performance even with the resource limits applied Part of the Gatling load test execution report. The result shows good performance even with the resource limits applied

And that is the end of my Django + ORM version for the rinha!

Conclusion

So is gevent bad and is it better to just use sync? Not exactly. What I can say is that in a scenario of very low latency, fast external requests, high load and limited resources, it is better to give up “async” to get connection reuse. The overhead of opening a connection per request was the determining factor for the API not passing the challenge’s load tests. I believe that using a library that implements a connection pool in Django would make it possible to use gevent for this case.

Also, the biggest lessons were:

  • Persistent connections and the gevent worker do not work together. You have to choose one or the other;
  • Using gunicorn, if you change the worker type from sync to gevent, run tests to choose the worker connections value. A value far above the ideal hurts performance quite a bit;
  • You cannot solve the problem without knowing what the problem is. In other words, use tools to measure and enable an effective analysis. I would not have gotten far without docker stats, pgAdmin and py-spy. And if I had spent more time on instrumentation I could have gone further and wasted less time on trial and error.

References

← All posts