Martin Kroeker
7d47f0a82d
Merge pull request #1978 from danielgindi/feature/msvc_cmake
Better support for MSVC/Windows in CMake (v0.3.x)
7 years ago
Martin Kroeker
a529c71a74
Merge pull request #1962 from brada4/r
Modrenize R benchmarks slightly
7 years ago
Martin Kroeker
89b60dab8a
Merge pull request #1987 from martin-frbg/issue1961
Change ARMV8 target with BINARY=32 to ARMV7 automatically
7 years ago
Martin Kroeker
58dd7e4501
Change ARMV8 target to ARMV7 for BINARY=32
7 years ago
Martin Kroeker
36b844af88
Change ARMV8 target to ARMV7 when BINARY32 is set
fixes #1961
7 years ago
Martin Kroeker
3f7bb87a2a
Merge pull request #1971 from martin-frbg/trsm-threshold
Shift transition to multithreading towards larger matrix sizes
7 years ago
Martin Kroeker
8533aca964
Avoid penalizing tall skinny matrices
7 years ago
Martin Kroeker
16494cb7c4
Merge pull request #1980 from martin-frbg/issue1979
Report SkylakeX as Haswell if compiler does not support AVX512
7 years ago
Martin Kroeker
b56b34a75c
Syntax fix
7 years ago
Martin Kroeker
21eda8b577
Report SkylakeX as Haswell if compiler does not support AVX512
... or make was invoked with NO_AVX512=1
7 years ago
Daniel Cohen Gindi
24288803b3
Adjust test script for correct deployment
7 years ago
Martin Kroeker
f0d834b824
Use VERSION_LESS for comparisons involving software version numbers
7 years ago
Daniel Cohen Gindi
63bbd7b0d7
Better support for MSVC/Windows in CMake
7 years ago
Martin Kroeker
010d59bfee
Merge pull request #1973 from martin-frbg/issue1464
Increase Zen SWITCH_RATIO to 16
7 years ago
Martin Kroeker
83b5c6b92d
Fix compilation with NO_AVX=1 set
fixes #1974
7 years ago
Martin Kroeker
bbfdd6c0fe
Increase Zen SWITCH_RATIO to 16
following GEMM benchmarks on Ryzen2700X. For #1464
7 years ago
Martin Kroeker
cda81cfae0
Shift transition to multithreading towards larger matrix sizes
See #1886 and JuliaRobotics issue 500. trsm benchmarks on Haswell and Zen showed that with these values performance is roughly doubled for matrix sizes between 8x8 and 14x14, and still 10 to 20 percent better near the new cutoff at 32x32.
7 years ago
Martin Kroeker
32b0f1168e
Fix declaration of input arguments in the Sandybridge GER microkernels ( #1967 )
* Tag arguments 0 and 1 as both input and output
7 years ago
Martin Kroeker
b495e54310
Fix declaration of input arguments in the x86_64 SCAL microkernels ( #1966 )
* Tag arguments 0 and 1 as both input and output (see #1964 )
7 years ago
Martin Kroeker
d5e6940253
Fix declaration of input arguments in the x86_64 microkernels for DOT and AXPY ( #1965 )
* Tag operands 0 and 1 as both input and output
For #1964 (basically a continuation of coding problems first seen in #1292 )
7 years ago
Martin Kroeker
24e697eadb
Merge pull request #1970 from quickwritereader/develop
crot fix
7 years ago
Martin Kroeker
3e9fd6359d
Bump xcode version to 10.1 to make sure it handles AVX512
7 years ago
Ubuntu
43a4572038
crot fix
7 years ago
Martin Kroeker
256eb588bb
Merge pull request #1963 from quickwritereader/develop
Blas1 single missing kernels implemented with vector builtins
7 years ago
Abdelrauf
a034e65512
Merge branch 'develop' into develop
7 years ago
Ubuntu
8c3386be87
Added missing Blas1 single fp {saxpy, caxpy, cdot, crot(refactored version of srot),isamax ,isamin, icamax, icamin},
Fixed idamin,icamin choosing the first occurance index of equal minimals
7 years ago
Andrew
3e601bd419
disable NaN checks before BLAS calls dgemm.R
7 years ago
Andrew
478d3c4569
disable NaN checks before BLAS calls deig.R (shorten matrix def)
7 years ago
Andrew
3afceb6c2a
disable NaN checks before BLAS calls deig.R
7 years ago
Andrew
7af8b21dbb
disable NaN checks before BLAS calls dsolve.R (shorter formula)
7 years ago
Martin Kroeker
1e3ada6db4
Merge pull request #1960 from cnjsdfcy/Hygon
Add support for Hygon Dhyana
7 years ago
Andrew
2777a7f506
disable NaN checks before BLAS calls dsolve.R (shorter config part)
7 years ago
Andrew
b70fd23836
disable NaN checks before BLAS calls dsolve.R
7 years ago
Andrew
def0385caa
init
7 years ago
caiyu
29dc72889f
Add support for Hygon Dhyana
7 years ago
Martin Kroeker
dbc9a060ef
Fix missing braces in support_av() call
7 years ago
Martin Kroeker
00401489c2
Fix missing braces in support_avx()
7 years ago
Martin Kroeker
21c0f2af7b
Merge pull request #1957 from martin-frbg/issue1954
Move TLS key deletion to openblas_quit
7 years ago
Martin Kroeker
ad2c386d6a
Move TLS key deletion to openblas_quit
fixes #1954 (as suggested by thrasibule in that issue)
7 years ago
Martin Kroeker
8d99dba86b
Merge pull request #1949 from martin-frbg/issue1947
Query AVX2 and AVX512VL support when selecting x86 kernels
7 years ago
Martin Kroeker
1650311246
Bump xcode to 8.3
7 years ago
Martin Kroeker
cf5d48e833
Update OSX environment to Sierra
as homebrew seems to have dropped support for El Capitan in their gcc packages
7 years ago
Martin Kroeker
191677b902
Add travis_wait to the OSX brew install phase
7 years ago
Martin Kroeker
31ed19e8b9
Add message for SkylakeX and KNL fallbacks to Haswell
7 years ago
Martin Kroeker
e1574fa2b4
Add xcr0 (os support) check
7 years ago
Martin Kroeker
68eb3146ce
Add xcr0 (os support) check
7 years ago
Martin Kroeker
0afaae4b23
Query AVX2 and AVX512VL capability in x86 cpu detection
7 years ago
Martin Kroeker
ae1d1f74f7
Query AVX2 and AVX512 capability for runtime cpu selection
7 years ago
Martin Kroeker
ed01f4932a
Merge pull request #1946 from martin-frbg/issue1908
More fixes for cross-compiling ARM64 targets
7 years ago
Martin Kroeker
802f0dbde1
More fixes for cross-compiling ARM64 targets
Fixed core naming for DYNAMIC_ARCH. Corrected GEMM_DEFAULT entries and added SYMV_P. Replaced outdated VULCAN define for ThunderX2T99 with ARMV8 to get basic definitions back. For issue #1908
7 years ago