Arjan van de Ven
cdc668d82b
Add a "sgemm direct" mode for small matrixes
OpenBLAS has a fancy algorithm for copying the input data while laying
it out in a more CPU friendly memory layout.
This is great for large matrixes; the cost of the copy is easily
ammortized by the gains from the better memory layout.
But for small matrixes (on CPUs that can do efficient unaligned loads) this
copy can be a net loss.
This patch adds (for SKYLAKEX initially) a "sgemm direct" mode, that bypasses
the whole copy machinary for ALPHA=1/BETA=0/... standard arguments,
for small matrixes only.
What is small? For the non-threaded case this has been measured to be
in the M*N*K = 28 * 512 * 512 range, while in the threaded case it's
less, around M*N*K = 1 * 512 * 512
7 years ago
Martin Kroeker
87718807f0
Merge pull request #1910 from martin-frbg/issue1909
Fix for DYNAMIC_ARCH builds made on a AVX512-capable host
7 years ago
Martin Kroeker
51aec8e96b
make sure the added march=skylake-avx512 does not cause problems on Windows
7 years ago
Martin Kroeker
06f7d78d70
Add -march=skylake-avx512 to SkylakeX part of DYNAMIC_ARCH builds
7 years ago
Martin Kroeker
38cc638591
Avoid adding blanket march=skylake-avx512 to dynamic_arch builds
7 years ago
Martin Kroeker
0bf6d74e5f
Fix typo in previous commit for arm dynamic arch
7 years ago
Martin Kroeker
133c278ee5
Add DYNAMIC_CORE list for ARM64
cf #1908
7 years ago
Martin Kroeker
2b355592e3
Make sure to use the arm version of dynamic.c in ARM64 DYNAMIC_ARCH
cf. #1908
7 years ago
Martin Kroeker
ff3eb1d474
Merge pull request #1904 from martin-frbg/issue1870
Fix cmake parsing of GEMM kernels for ARMV8
7 years ago
Martin Kroeker
0b09516678
Fix missing parameter in popen call
7 years ago
Martin Kroeker
7639f2e1f0
Rewrite the conditional for OSX to fix cmake parsing on others
The Makefile variable parser in utils.cmake currently does not handle conditionals. Having the definitions for non-OSX last will at least make cmake builds work again on non-OSX platforms.
7 years ago
Martin Kroeker
2fc712469d
Avoid creating spurious non-suffixed c/zgemm_kernels
Plain cgemm_kernel and zgemm_kernel are not used anywhere, only cgemm_kernel_b etc.
Needlessly building them (without any define like NN, CN, etc.) just happened to work on most platforms, but not on arm64. See #1870
7 years ago
Martin Kroeker
6ba30e270d
Fix typo that broke CNRM2 on ARMV8 since 0.3.0
must have happened in my #1449
7 years ago
Martin Kroeker
bf23518e36
Merge pull request #1903 from rengolin/armv8
Fix two mistakes on Arm64 builds
7 years ago
Renato Golin
31a490ea88
Fix two mistakes on Arm64 builds
* Falkor is an ARMv8.0 with ARMv8.1 features, and chosing armv8.1-a for
march generates instructions it cannot cope with. Reverting it back
to armv8-a.
* ThunderX2's build was left with a #define VULCAN, which made it miss
the right compiler flags in Makefile.arm64, although it did create
the right library in the end.
7 years ago
Martin Kroeker
701ea88347
Use p2align instead of align for OSX compatibility
fixes #1902
7 years ago
Martin Kroeker
721c56c224
Merge pull request #1899 from brada4/fbsd12
Add mutually supported architecture mappings for FreeBSD12 ports
7 years ago
Martin Kroeker
c5f8aeff2d
Merge branch 'develop' into fbsd12
7 years ago
Martin Kroeker
8278cbe7f8
Merge pull request #1894 from pkubaj/patch-2
Use correct ARCH name on BSD powerpc64
7 years ago
Martin Kroeker
ea6d1b96bd
Update Makefile.system
7 years ago
Martin Kroeker
360374be62
Update with the changes from 0.3.4
7 years ago
Martin Kroeker
f5acaad8f0
Increment version to 0.3.5.dev
7 years ago
Martin Kroeker
93fa6b7b76
Increment version to 0.3.5.dev
7 years ago
Martin Kroeker
b028960aba
Merge branch 'release-0.3.0' into develop
7 years ago
Martin Kroeker
3c9e3faedb
fixup BSD naming of powerpc arch
7 years ago
Andrew
44c81fd135
oops
7 years ago
Andrew
26b3710485
Add architecture mappings for FreeBSD12
7 years ago
Andrew
84e614d0fd
init
7 years ago
Martin Kroeker
dceff5542c
Handle Android environments that identify as Linux ( #1898 )
* Handle Android environments that identify as Linux
termux terminal emulator does this, causing build failures through missed defines in common.h
7 years ago
Martin Kroeker
6c7b691083
Really revert xDOT changes from 1832
neglected to rebase #1892 on merging
7 years ago
Martin Kroeker
5f4c550c27
Merge pull request #1892 from martin-frbg/mipsdot
revert MIPS64 xDOT kernel changes from #1832
7 years ago
pkubaj
731b2722ba
Fix build on POWER, remove DragonFly, add NetBSD
__asm is complete on its own
DBSD developers state they will only support amd64, but NetBSD supports POWER.
7 years ago
pkubaj
f85ce54d4a
Use correct Makefile on powerpc64
FreeBSD uses powerpc64 name for POWER architecture. Use correct Makefile for this platform.
7 years ago
Andrew
2601cd58ab
remove surplus locking code , only enabled w x86, disabled or never enabled on all others
7 years ago
Martin Kroeker
95a5542e3c
Revert DOT kernel changes from #1834
as the failures seen on Loongson3A appear to be limited to DSDOT/SDSDOT (i.e. my hackish "fix" from #1684 )
7 years ago
Martin Kroeker
7a2e1bc804
Use generic kernel for DSDOT/SDSDOT
as discussed in #1834
7 years ago
Martin Kroeker
35653e38b3
Merge pull request #1834 from fengrl/develop
register push/pop command change
7 years ago
Martin Kroeker
71e25ae42f
Merge pull request #1890 from martin-frbg/issue1889
Include version number in openblas_get_config output
7 years ago
Martin Kroeker
97d7298973
call it OpenBLAS not just version
7 years ago
Martin Kroeker
de0d0ed52f
Improve formatting of config output
7 years ago
Martin Kroeker
081ceb3e02
Propagate version number for openblas_get_config
7 years ago
Martin Kroeker
a29ec458c2
propagate verison number for openblas_config_version
7 years ago
Martin Kroeker
816775e309
Add version information to openblas_get_config output
7 years ago
Martin Kroeker
b6363f4539
Merge pull request #1885 from brada4/freebsd
Fix freebsd clang compilation of skylakex
7 years ago
Andrew
19c4bdd8b3
Add return value so that freebsd system clang does not err out
7 years ago
Andrew
f049a4c84f
init
7 years ago
Martin Kroeker
f72fdf525c
Merge pull request #1875 from martin-frbg/issue1851
Serialize accesses to parallelized level3 functions from multiple cal…
7 years ago
Martin Kroeker
5393759a98
Merge pull request #1869 from martin-frbg/axpy0
Handle special case INCX=0,INCY=0 in the axpy interface
7 years ago
Martin Kroeker
5cf18e2875
Merge pull request #1878 from kiwifb/PGI_f_check
Correct link flags for PGI compiler.
7 years ago
Martin Kroeker
910050985a
Merge pull request #1876 from rengolin/armv8-cleanup
Simplifying ARMv8 build parameters
7 years ago