In its estimation of the 10 most important emerging technologies for the year 2010, Technology Review mentions a programming language.
Yes, a language.
When we started Ateji, the common wisdom was that language was an irrelevant artifact of the programming process. I'm glad to see that we were on the right track, and that the importance of language is now acknowledged even in the most prestigious magazines.
We're preparing an offering for cloud programming that will let the same source code run indifferently on desktop PCs, multi-core servers, clusters, or in the cloud, while remaining compatible with existing tools and practice. This takes the form of a language extension, and wouldn't have been possible without language technology and clever language design.
Tuesday, April 27, 2010
A *language* in the top-10 technologies
Wednesday, March 31, 2010
Robin Milner
Professor Robin Milner, one of the pioneers of computer science, passed away on March 20: http://www.timesonline.co.uk/tol/comment/obituaries/article7081867.ece.
He will be remembered, among many other achievements, as the inventor of pi-calculus, the theoretical foundation behind Ateji Parallel Extensions.
Though I never had a chance to actually meet him, he inspired my career all along, as a researcher and as an entrepreneur.
Saturday, March 27, 2010
Ateji PX vs Java's ParallelArray
The ParallelArray API is a proposal for integrating data parallelism into the Java language. Here is an overview by Brian Goetz: http://www.ibm.com/developerworks/java/library/j-jtp11137.html
The canonical example for ParallelArray shows how to compute an average of student grades in parallel. Note the declaration of students as a ParallelArray rather than an array, the introduction of a filter 'isSenior' and a selector 'selectGpa', and the use of an ExecutorService 'fjPool'.
ParallelArray
double bestGpa = students.withFilter(isSenior)
.withMapping(selectGpa)
.max();
static final Ops.Predicate
public boolean op(Student s) {
return s.graduationYear == Student.THIS_YEAR;
}
};
static final Ops.ObjectToDouble
public double op(Student student) {
return student.gpa;
}
};
Closures may help in making this code less verbose. Here is an example based on the BGGA proposal:
ParallelArray students = new ParallelArray(fjPool, data);
double bestGpa = students.withFilter({Student s => (s.graduationYear == THIS_YEAR) })
.withMapping({ Student s => s.gpa })
.max();
However, what we're interested in with this use case is expressing so-called "reduction operations" over arrays using generators, filters and selectors, and run them in parallel. Making Java a functional language is an interesting but different topic.
Here is the syntax used by Ateji Parallel Extensions for expressing reduction operations (we actually call them "comprehensions", as in "list comprehension"). This syntax should look familiar, as it inspired by the set notation used in high school mathematics.
Student[] students = ...;
double bestGpa = `+ for{ s.gpa | Student s : students, if s.graduationYear == THIS_YEAR };
Note that
students remains a classical Java array. This code is sequential, it becomes parallel when a parallel operator is inserted right after the for keyword:Student[] students = ...;
double bestGpa = `+ for||{ s.gpa | Student s : students, if s.graduationYear == THIS_YEAR };
The parallel operator on reduction expressions provides almost linear speedup on multi-core processors, as soon as the amount of computation for each array element is more than a couple of additions.
Quiz:
Using the comprehension notation, how would you express the number of students graduating this year ? The average of their grades ?
Tuesday, March 23, 2010
Matrix multiplication - 12.5x speedup with a single "||"
Agreed, matrix multiplication is not a very sexy piece of code, but it serves as a standard example and benchmark for data-parallel computations (code that applies the same operation to many elements in parallel).
The matrix multiplication white-paper is available from http://www.ateji.com/multicore/whitepapers.html. It shows how we achieved a 12.5x speedup on a 16-core server, simply adding one single "||" operator to an existing sequential Java code. Raw performance is pretty good as well, on par with linear algebra libraries.
Here is the parallel code. Note the "||" operator right after the first 'for' keyword, this is the only difference between sequential and parallel version of the code.
for||(int i : I) {
for(int j : J) {
for(int k : K) {
C[i][j] += A[i][k] * B[k][j];
}
}
}
Performance is pretty impressive, on par with dedicated linear algebra librairies:
The part that I find really interesting is the comparison with the same algorithm using plain Java threads. Even if you have a general knowledge about threads, you need to see actual code before you can imagine the amount of small details that need to be taken into account.
They include adding many final keywords, copying local variables, computing indices, managing InterruptedException. 27 lines vs. 7 lines. And we haven't even returned values or thrown exception from within threads! The problem is not so much verbosity itself, but the fact that programmer's intent gets hidden behind a lot of irrelevant details.
Enjoyed the article?
Share your interest by voting up this article on social sites!
Tuesday, November 24, 2009
100-cores by next year
Once again, hardware is far ahead of software. Tilera has announced a 100-cores processor for 2010.
Unlike standard multi-core processors, Tilera's TILE-Gx is architectured around a 2D grid network rather than a single shared bus. This is a way to jump over the "memory wall", and feed enough data to keep all cores busy.
This design provides a lot of raw computing power, but also better efficiency (more computing power per watt).
But this beast is supposed to be coded in standard C/C++. It already requires black magic in order to write a 2-threads program that behaves as expected, what about 100's of threads ?
The TILE-Gx is a perfect match for Ateji Parallel Extensions : data parallelism handles large scientific computations task parallelism handles server-like applications, and message-passing leverages the hardware's packet network interconnection mechanism. High-performance code can be arranged in a data-flow or streaming architecture, reducing accesses to shared memory.
Wednesday, November 18, 2009
Session Evaluation
I just received my session evaluation from the TSS Java Symposium Europe : an impressive 4.58/5.0 with the comment "Great Session !".
I am less proud of the speaker evaluation, at 4.21/5.0. If you attended the session, I'd be happy to hear from you about what could be improved.
Sunday, November 1, 2009
Parallelism at the language level - Part 1: Hello World
The major contribution of Ateji Parallel Extensions is to add parallelism at the language level.
What does this change? Today's mainstream programming languages have been designed with sequential processing in mind, they simply have no idea about what is parallelism. Consider how you'd run two tasks in parallel in Java:
void run() {
println("Hello"); // print Hello in the other thread
}
}.start();
println("World"); // print World in this thread
otherThread.join(); // wait until code1 has terminated
Not to mention how unreadable and unmaintainable this code is, you'll notice that there is a fair amount of black magic involved here: just because you called a method whose name happens to be
start(), the whole behaviour of your program has changed. But the compiler is not aware of this change, it thinks it is just calling an ordinary library method.With Ateji Parallel Extensions, two tasks are run in parallel by composing them using the || operator :
How could it be simpler?
Not only is this much more concise and understandable, it also makes it easier for the developer to "think" parallel and to catch potential errors early.
And since the very idea of parallelism is present in the language, the compiler is able to understand the actual meaning of the code and to perform tricks such as high-level code optimization or better verification.
Read more on parallelism at the language level.
