Fixes#2496 and includes a non-regression test. Since the introduction
of fast_dot (and NumPy 1.6), dot products between arrays of various
memory layouts have become fast enough to no longer require storing coef_
in Fortran order. The only noticeably slower operation is multiplying a
CSR matrix with a Fortran-ordered array, but then building the CSR matrix
in the first place is more likely to be the bottleneck.