Showing posts with label Programming. Show all posts
Showing posts with label Programming. Show all posts

Manual profiling of python code (especially if you use celery)




English: Official logo of the Université libre...
English: Official logo of the Université libre de Bruxelles (Photo credit: Wikipedia)

Profiling tools like cprofile provide hardly understandable output for my celery based app. So I used a tool called manual profiling by Pierre de Buyl from University of Toronto - Université Libre de Bruxelles (many thanks Pierre for your help :) )
It's a great tool. You import it in the source code that you want to profile:

from manual_profiler import Profiler
import warnings
I import warnings too, so that the print commands do not get lost in celery's arcanes. Then, in my code, I create an instance of the profiler, and register all the functions that I want to monitor: 

  # profiling
pr=Profiler()
correctErrors__=pr.register_function(correctErrors)
findSpellErrors__=pr.register_function(findSpellErrors)

As you can see, the register_function method rewrites the function to add the timing features to it. So I call the rewritten functions in my code.
Then, at the end of my code, I add a command to print the results:

    p.display()

With this configuration I obtain something like this:
[2012-11-27 22:24:35,937: WARNING/PoolWorker-1] Profiling information
--------------------------------------------------------------------------------
Name # of calls total time time per call
--------------------------------------------------------------------------------
[2012-11-27 22:24:35,937: WARNING/PoolWorker-1] correctErrors 41 64.102476 s 1.563475 s
[2012-11-27 22:24:35,938: WARNING/PoolWorker-1] findSpellErrors 1 61.000884 s 61.000884 s
Note that I improved the code a little bit to sort the entries by total time. I added the following line into the display function:

self.timers = sorted(self.timers, key=lambda t:t[1]._time, reverse=True)

This profiler gives information at functional level. Now, that I know that correctErrors is the function I need to improve, I use create the following class to monitor intra function events:
class ProfilingTimers():
""" define a few timers for manual profiling"""
def __init__(self):
"""Instancing the class prints the current time"""
self.start=datetime.now()
self.last=datetime.now()
self.elapsed=self.last - self.start
print strftime("Start time: "+str(self.start)) def elapsedSinceLast(self, eventName="N/A"):
""" Calling this method displays a the elapsed time since the last event"""
self.elapsed=datetime.now() - self.last
self.last=datetime.now()
print strftime("End of "+eventName+" - Time elapsed since last event: "+str(self.elapsed.days)+" days "+str(self.elapsed.seconds)+' s '+str(self.elapsed.microseconds/1000)+" ms")

In the code  at the start of the section of code to be monitored, I add :
p= ProfilingTimers()

And in a few places afterwards:

   p.elapsedSinceLast('image processing')

This command prints the elapsed time since the command was last invoked or since the ProfilingTimers instance was created. In the printed message, it includes the eventname that you have specified (here 'image processing'). With this method I can accurately assess with section of my function's code takes the most time.

With this great method, I could dig into my problem and after a few improvements... I obtained a 3200 times faster code execution!!!!

[2012-11-27 22:51:03,779: WARNING/PoolWorker-2] Profiling information
--------------------------------------------------------------------------------
Name # of calls total time time per call
--------------------------------------------------------------------------------
[2012-11-27 22:51:03,780: WARNING/PoolWorker-2] correctErrors 41 0.027236 s 0.000664 s
[2012-11-27 22:51:03,780: WARNING/PoolWorker-2] findSpellErrors 1 1.653973 s 1.653973 s


The poor man's method would be much less practical... It would involve putting in a few places the following command :

print strftime("time: %H:%M:%S", gmtime())
 
 
Share:
Read More

Profiling python code with celery

In this article I explain how you can profile python code with celery, and why I find this solution disappointing. I propose a better solution in the conclusion. Have fun !
Head of celery, sold as a vegetable. Usually o...
Head of celery, sold as a vegetable. Usually only the stalks are eaten. (Photo credit: Wikipedia)
If you are in a hurry, you can jump to the conclusion about python profiling tools at the end of this post. Otherwise, you will find below the summary of the tests I performed with the most common python profiling software.

First I install celerymon:

pip install celerymon

Then to run my celery powered module, I add the -E option:

celery -A mypackage.mymodule worker --loglevel=info -E

At this point, the events monitored by celerymon are available at:

firefox http://localhost:8989/

Celerymon displays few events, so it is not adapted for code profiling.

For code profiling, I try using cprofile. So, to launch celery now I use a different command

sudo apt-get install python-profiler
python -m cProfile -o test-`date +%Y-%m-%d-%T`.prof /home/toto/virtualenv_1/bin/celery -A mypackage.mymodule worker --loglevel=info -E

Alternatively, it's possible to modify the python code to include cProfile directives (but I have not yet managed to collect the output in a file):

import cProfile
cProfile.run('foo()','filename.prof')

The profiling data is pure text, and hard to manipulate, even with pstats. So, I use visualization tools.

### KCACHEGRIND ###
kcachegrind is probably the best tool today to analyse profiling data.

sudo apt-get install kcachegrind
easy_install pyprof2calltree
pyprof2calltree -i myfile.prof -o myfile.prof.grind
kcachegrind myfile.prof.grind

### RUNSNAKERUN ###
runsnakerun is a more recent tool, with more limited functionalities

pip install SquareMap RunSnakeRun
runsnake OpenGLContext.profile

With this tool, you have a nice display of all the calls, the cumulative time spent per function, etc.
If you have the following problem, reinstall wxpython:
    from squaremap import squaremap
File "/home/toto/virtualenv_1/lib/python2.6/site-packages/squaremap/squaremap.py", line 3, in
import wx.lib.newevent
ImportError: No module named lib.newevent
(virtualenv_1)toto:~/virtualenv_1/djangoProj_1$ pip install wxpython

### CONCLUSION ###

None of the above tools was handy for my app. They did not allow me to see clearly what lines of my code where taking the most time. Almost all the time seemed to be spent in Kombu module which is used by AMPQ. See my next post about manual profiling to see how I progressed nevertheless and managed to divide by 3000 the time spent in my most time consuming function!





Share:
Read More
,

How to replicate / export a virtualenv from a machine to another ?

Sudo
Sudo (Photo credit: Scelus' Comix)
If you code in Python, you probably use virtualenv to maintain a clean setup and avoid problems with different versions of pythons. But sometimes, you need to replicate / export a virtualenv from a machine to another. This post will explain you exactly how to do this.

First, use the following commands on the origin machine:

source ./bin/activate
pip freeze > requirements.txt

Then, copy your source code files from the virtualenv folder (called virtualenv_1) on the target machine. Do NOT copy the virtualenv files in folders such as bin, lib etc. We will eventually create these files with virtualenv on the new system:

sudo apt-get install python-virtualenv
sudo  virtualenv virtualenv_1

Then install all the required packages:

cd virtualenv_1
source bin/activate
sudo pip install -r requirements.txt

That's it! If you have liked this article, please put a link to it on your Google+  / facebook profile are any other kind of web site: it will help others find it in the search engines. Thanks!




Share:
Read More
,

Celery and warning messages (print) to stdout

In this post, I will explain you how to use Celery and print warning messages to stdout.
Cross section of celery stalk, showing vascula...
Cross section of celery stalk, showing vascular bundles, which include both phloem and xylem. (Photo credit: Wikipedia)
If you use celery, you might be frustrated by the absence of the 'print' assertions that you put in your code for debugging. There is a very simple way to solve this point and see all print. As an added benefit, it will permit you to display nice warning in a python way :).

The solution is very simple!  Just import the warning package in the module that you call with celery:
 
import warnings 
 
 
Share:
Read More

why "from toto Import *" is a bad idea in Python programs.

In this post, I will explain you why the good practice in Python programs is to carefully import only what is necessary to run your program. Let's start with an example. Sometimes, I feel lazy and use:

from toto import *

This is a bad idea, as sometimes the name of the things that you import will conflict with the name of other things in your programs. And this can be a nightmare to debug. So, try not to be lazy and rather define explicitly what you want to import with:

import toto
or
from toto import titi, tata

Share:
Read More
,

A simple DeTeX function in python - LaTeX to text

The LaTeX logo, typeset with LaTeX
The LaTeX logo, typeset with LaTeX (Photo credit: Wikipedia)
I have implemented a simple DeTeX function in Python. I provide this function below, as is and without any guarantee. If you run it, and it should change the example LaTeX text into "simple" text thanks the detex() function defined in the code.

It's a quick and dirty approach: I did not try to implement the full LaTeX syntax. I just applied a few regexps to strip the commands of the text. Feedback will be appreciated in the comment form below :)

Take care to the "backslash plague" as explained in http://docs.python.org/2/howto/regex.html".

#!/usr/bin/python
# -*- coding: UTF-8 -*-

import re

testMode=False

def applyRegexps(text, listRegExp):
""" Applies successively many regexps to a text"""
if testMode:
print '\n'.join(listRegExp)
# apply all the rules in the ruleset
for element in listRegExp:
left = element['left']
right = element['right']
r=re.compile(left)
text=r.sub(right,text)
return text

"""
_ _ ____
__| | ___| |_ _____ __/ /\ \
/ _` |/ _ \ __/ _ \ \/ / | | |
| (_| | __/ || __/> <| | | |
\__,_|\___|\__\___/_/\_\ | | |
\_\/_/
"""

def detex(latexText):
"""Transform a latex text into a simple text"""
# initialization
regexps=[]
text=latexText
# remove all the contents of the header, ie everything before the first occurence of "\begin{document}"
text = re.sub(r"(?s).*?(\\begin\{document\})", "", text, 1)

# remove comments
regexps.append({r'left':r'([^\\])%.*', 'right':r'\1'})
text= applyRegexps(text, regexps)
regexps=[]

# - replace some LaTeX commands by the contents inside curly rackets
to_reduce = [r'\\emph', r'\\textbf', r'\\textit', r'\\text', r'\\IEEEauthorblockA', r'\\IEEEauthorblockN', r'\\author', r'\\caption',r'\\author',r'\\thanks']
for tag in to_reduce:
regexps.append({'left':tag+r'\{([^\}\{]*)\}', 'right':r'\1'})
text= applyRegexps(text, regexps)
regexps=[]
"""
_ _ _ _ _ _
| |__ (_) __ _| (_) __ _| |__ | |_
| '_ \| |/ _` | | |/ _` | '_ \| __|
| | | | | (_| | | | (_| | | | | |_
|_| |_|_|\__, |_|_|\__, |_| |_|\__|
|___/ |___/
"""
# - replace some LaTeX commands by the contents inside curly brackets and highlight these contents
to_highlight = [r'\\part[\*]*', r'\\chapter[\*]*', r'\\section[\*]*', r'\\subsection[\*]*', r'\\subsubsection[\*]*', r'\\paragraph[\*]*'];
# highlightment pattern: #--content--#
for tag in to_highlight:
regexps.append({'left':tag+r'\{([^\}\{]*)\}','right':r'\n#--\1--#\n'})
# highlightment pattern: [content]
to_highlight = [r'\\title',r'\\author',r'\\thanks',r'\\cite', r'\\ref'];
for tag in to_highlight:
regexps.append({'left':tag+r'\{([^\}\{]*)\}','right':r'[\1]'})
text= applyRegexps(text, regexps)
regexps=[]

"""
_ __ ___ _ __ ___ _____ _____
| '__/ _ \ '_ ` _ \ / _ \ \ / / _ \
| | | __/ | | | | | (_) \ V / __/
|_| \___|_| |_| |_|\___/ \_/ \___|

"""
# remove LaTeX tags
# - remove completely some LaTeX commands that take arguments
to_remove = [r'\\maketitle',r'\\footnote', r'\\centering', r'\\IEEEpeerreviewmaketitle', r'\\includegraphics', r'\\IEEEauthorrefmark', r'\\label', r'\\begin', r'\\end', r'\\big', r'\\right', r'\\left', r'\\documentclass', r'\\usepackage', r'\\bibliographystyle', r'\\bibliography', r'\\cline', r'\\multicolumn']

# replace tag with options and argument by a single space
for tag in to_remove:
regexps.append({'left':tag+r'(\[[^\]]*\])*(\{[^\}\{]*\})*', 'right':r' '})
#regexps.append({'left':tag+r'\{[^\}\{]*\}\[[^\]\[]*\]', 'right':r' '})
text= applyRegexps(text, regexps)
regexps=[]

"""
_
_ __ ___ _ __ | | __ _ ___ ___
| '__/ _ \ '_ \| |/ _` |/ __/ _ \
| | | __/ |_) | | (_| | (_| __/
|_| \___| .__/|_|\__,_|\___\___|
|_|
"""

# - replace some LaTeX commands by the contents inside curly rackets
# replace some symbols by their ascii equivalent
# - common symbols
regexps.append({'left':r'\\eg(\{\})* *','right':r'e.g., '})
regexps.append({'left':r'\\ldots','right':r'...'})
regexps.append({'left':r'\\Rightarrow','right':r'=>'})
regexps.append({'left':r'\\rightarrow','right':r'->'})
regexps.append({'left':r'\\le','right':r'<='})
regexps.append({'left':r'\\ge','right':r'>'})
regexps.append({'left':r'\\_','right':r'_'})
regexps.append({'left':r'\\\\','right':r'\n'})
regexps.append({'left':r'~','right':r' '})
regexps.append({'left':r'\\&','right':r'&'})
regexps.append({'left':r'\\%','right':r'%'})
regexps.append({'left':r'([^\\])&','right':r'\1\t'})
regexps.append({'left':r'\\item','right':r'\t- '})
regexps.append({'left':r'\\\hline[ \t]*\\hline','right':r'============================================='})
regexps.append({'left':r'[ \t]*\\hline','right':r'_____________________________________________'})
# - special letters
regexps.append({'left':r'\\\'{?\{e\}}?','right':r'é'})
regexps.append({'left':r'\\`{?\{a\}}?','right':r'à'})
regexps.append({'left':r'\\\'{?\{o\}}?','right':r'ó'})
regexps.append({'left':r'\\\'{?\{a\}}?','right':r'á'})
# keep untouched the contents of the equations
regexps.append({'left':r'\$(.)\$', 'right':r'\1'})
regexps.append({'left':r'\$([^\$]*)\$', 'right':r'\1'})
# remove the equation symbols ($)
regexps.append({'left':r'([^\\])\$', 'right':r'\1'})
# correct spacing problems
regexps.append({'left':r' +,','right':r','})
regexps.append({'left':r' +','right':r' '})
regexps.append({'left':r' +\)','right':r'\)'})
regexps.append({'left':r'\( +','right':r'\('})
regexps.append({'left':r' +\.','right':r'\.'})
# remove lonely curly brackets
regexps.append({'left':r'^([^\{]*)\}', 'right':r'\1'})
regexps.append({'left':r'([^\\])\{([^\}]*)\}','right':r'\1\2'})
regexps.append({'left':r'\\\{','right':r'\{'})
regexps.append({'left':r'\\\}','right':r'\}'})
# strip white space characters at end of line
regexps.append({'left':r'[ \t]*\n','right':r'\n'})
# remove consecutive blank lines
regexps.append({'left':r'([ \t]*\n){3,}','right':r'\n'})
# apply all those regexps
text= applyRegexps(text, regexps)
regexps=[]
# return the modified text
return text

"""
_
_ __ ___ __ _(_)_ __
| '_ ` _ \ / _` | | '_ \
| | | | | | (_| | | | | |
|_| |_| |_|\__,_|_|_| |_|

"""
def main():
""" Just for debugging"""
#print "defining the test text\n"
latexText=r"""
% This paper can be formatted using the peerreviewca
% (instead of conference) mode.
\documentclass[twocolumn,a4paper]{article}
%\documentclass[peerreviewca]{IEEEtran}
% correct bad hyphenation here
\hyphenation{op-ti-cal net-works semi-con-duc-tor IEEEtran pri-va-cy Au-tho-ri-za-tion}
% package for printing the date and time (version)
\usepackage{time}
\begin{document}
\title{Next Generation Networks}
\author{Tot titi\thanks{Network and Security -- test company -- toto@ieee.org}}
\maketitle
\begin{abstract}\footnote{Version : \today ; \now}
lorem ipsum(\ldots)\end{abstract}
\emph{Keywords: IP Multimedia Subsystem, Quality of Service}
\section{Introduction} \label{sect:introduction}
lorem ipsum(\ldots) \% of the world population. \cite{TISPAN2006a}. \footnote{Bearer Independent Call Control protocol}.
\hline
\section{Protocols used in IMS} \label{sect:protocols}
lorem ipsum(\ldots) \cite{rfc2327, rfc3264}.
\subsection{Authentication, Authorization, and Accounting} \label{sect:protocols_aaa}
lorem ipsum(\ldots)
\subsubsection{Additional protocols} \label{sect:protocols_additional}
lorem ipsum(\ldots)
\begin{table}
\begin{center}
\begin{tabular}{|c|c|c|}
\hline
\textbf{Capability} & \textbf{UE} & \textbf{GGSN} \\ \hline
\emph{DiffServ Edge Function} & Optional & Required \\ \hline
\emph{RSVP/IntServ} & Optional & Optional \\ \hline
\emph{IP Policy Enforcement Point} & Optional & Required \\ \hline
\end{tabular}
\caption{IP Bearer Services Manager capability in the UE and GGSN}
\label{tab_ue_ggsn}
\end{center}
\end{table}
The main transport layer functions are listed below:
\begin{my_itemize}
\item The \emph{Resource Control Enforcement Function} (RCEF) enforces policies under the control of the A-RACF. It opens and closes unidirectional filters called \emph{gates} or \emph{pinholes}, polices traffic and marks IP packets \cite{TISPAN2006c}.
\item The \emph{Border Gateway Function} (BGF) performs policy enforcement and Network Address Translation (NAT) functions under the control of the S-PDF. It operates on unidirectional flows related to a particular session (micro-flows) \cite{TISPAN2006c}.
\item The \emph{Layer 2 Termination Point} (L2TP) terminates the Layer 2 procedures of the access network \cite{TISPAN2006c}.
\end{my_itemize}
Their QoS capabilities are summarized in table \ref{tab_rcef_bgf} \cite{TISPAN2006c}.
The admission control usually follows a three step procedure:
\begin{my_enumerate}
\item Authorization of resources (\eg by the A-RACF)
\item Resource reservation (\eg by the BGF)
\item Resource commitment (\eg by the RCEF)
\end{my_enumerate}
\begin{figure}
\centering
\includegraphics[width=1.5in]{./pictures/RACS_functional_architecture}
\caption{RACS interaction with transfer functions}
\label{fig_RACS_functional_architecture}
\end{figure}
%\subsection{Example} \label{sect:qos_example}
% conference papers do not normally have an appendix
% use section* for acknowledgement
\section*{Acknowledgment}
% optional entry into table of contents (if used)
%\addcontentsline{toc}{section}{Acknowledgment}
lorem ipsum(\ldots)
\bibliographystyle{plain}
%\bibliographystyle{alpha}
\bibliography{./mabiblio}
\end{document}
"""
#print '\n'.join(diff)
text=detex(latexText)
print text


if __name__ == "__main__":
main()
Enjoy!  And feel free to comment below or to put a link to this article on your blog. Thanks!


Share:
Read More
, ,

A Django tutorial on Debian - part 4: creating the web administration portal

Django logo
Django logo (Photo credit: Wikipedia)
.Django creates automatically a web admin portal for our web app. We just have to edit the settings.py file and to uncomment "django.contrib.admin" in the INSTALLED_APPS setting.

    # Uncomment the next line to enable the admin:
'django.contrib.admin',
# Uncomment the next line to enable admin documentation:
'django.contrib.admindocs', 
 
We need to update the database to take into account the new apps. But first, we will check if we have the latest version of Django.

sudo easy_install --upgrade django
chmod +x manage.py
./manage.py syncdb
If we have a look to our django powered website, we will see that it is still empty, because we have not defined any url for the moment: 

./manage.py runserver
firefox http://127.0.0.1:8000/
Let's define the URLs by editing the urls.py file and uncommenting all lines related to the admin and admin documentation: 

# Uncomment the next two lines to enable the admin:
from django.contrib import admin
admin.autodiscover()

(...)

# Uncomment the admin/doc line below to enable admin documentation:
url(r'^admin/doc/', include('django.contrib.admindocs.urls')),

# Uncomment the next line to enable the admin:
url(r'^admin/', include(admin.site.urls)),
Now, our admin site is ready: 

firefox http://127.0.0.1:8000/admin/
If you do not remember the admin password that you setup in part 2 of this tutorial (the login is root), then create a new one.

The admin portal does not display our application textModif. We have to go to textModif folder and add a file admin.py there:

cd textModif
vim admin.py
We put the following lines in this file: 

from django.contrib import admin
from textModif.models import *

admin.site.register(TextTask)
admin.site.register(TextOperation)
admin.site.register(TextModification)
We restart the server to see the changes (kill the associated process if you have maintained the server): 

./manage.py runserver
firefox http://127.0.0.1:8000/admin
Now we can edit our data structure through the web administration interface. Django's official tutorial shows also how to customize very easily the admin interface. We will skip this part, as I would like to focus on our web app's frontend implementation.



Share:
Read More