Questions about opencl
#1
Hello,

and sorry for bad english, I have a few questions:
installed hashcat according to this manual

1. I can't use my first and third device at the same time??(which hashcat discovered)
    i have to choose one of them?
    if yes, why can't i use my third device?
3. My integrated video adapter( Intel Corporation UHD Graphics (rev 05)) cannot be used for brute-force with nvidia + cpu?
3. Is it possible to safely improve the brute-force speed?

Information about my system:
Code:
Linux pc 5.11.0-41-generic #45~20.04.1-Ubuntu SMP Wed Nov 10 10:20:10 UTC 2021 x86_64 x86_64 x86_64 GNU/Linux

nvidia-smi
Code:
Sun Dec  5 11:46:24 2021      +-----------------------------------------------------------------------------+ | NVIDIA-SMI 495.44      Driver Version: 495.44      CUDA Version: 11.5    | |-------------------------------+----------------------+----------------------+ | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC | | Fan  Temp  Perf  Pwr:Usage/Cap|        Memory-Usage | GPU-Util  Compute M. | |                              |                      |              MIG M. | |===============================+======================+======================| |  0  NVIDIA GeForce ...  Off  | 00000000:01:00.0 Off |                  N/A | | N/A  41C    P8    1W /  N/A |    396MiB /  3911MiB |      5%      Default | |                              |                      |                  N/A | +-------------------------------+----------------------+----------------------+                                                                               +-----------------------------------------------------------------------------+ | Processes:                                                                  | |  GPU  GI  CI        PID  Type  Process name                  GPU Memory | |        ID  ID                                                  Usage      | |=============================================================================| |    0  N/A  N/A      1030      G  /usr/lib/xorg/Xorg                45MiB | |    0  N/A  N/A      1656      G  /usr/lib/xorg/Xorg                156MiB | |    0  N/A  N/A      1830      G  /usr/bin/gnome-shell              41MiB | |    0  N/A  N/A      2119      G  /usr/lib/firefox/firefox          137MiB | |    0  N/A  N/A      2255      G  /usr/lib/firefox/firefox            1MiB | |    0  N/A  N/A      2483      G  /usr/lib/firefox/firefox            1MiB | |    0  N/A  N/A      3874      G  /usr/lib/firefox/firefox            1MiB | +-----------------------------------------------------------------------------+



clinfo
Code:
Number of platforms                              3   Platform Name                                  Portable Computing Language   Platform Vendor                                The pocl project   Platform Version                                OpenCL 1.2 pocl 1.4, None+Asserts, LLVM 9.0.1, RELOC, SLEEF, DISTRO, POCL_DEBUG   Platform Profile                                FULL_PROFILE   Platform Extensions                            cl_khr_icd   Platform Extensions function suffix            POCL   Platform Name                                  NVIDIA CUDA   Platform Vendor                                NVIDIA Corporation   Platform Version                                OpenCL 3.0 CUDA 11.5.100   Platform Profile                                FULL_PROFILE   Platform Extensions                            cl_khr_global_int32_base_atomics cl_khr_global_int32_extended_atomics cl_khr_local_int32_base_atomics cl_khr_local_int32_extended_atomics cl_khr_fp64 cl_khr_3d_image_writes cl_khr_byte_addressable_store cl_khr_icd cl_khr_gl_sharing cl_nv_compiler_options cl_nv_device_attribute_query cl_nv_pragma_unroll cl_nv_copy_opts cl_nv_create_buffer cl_khr_int64_base_atomics cl_khr_int64_extended_atomics cl_khr_device_uuid cl_khr_pci_bus_info   Platform Host timer resolution                  0ns   Platform Extensions function suffix            NV   Platform Name                                  Intel(R) CPU Runtime for OpenCL(TM) Applications   Platform Vendor                                Intel(R) Corporation   Platform Version                                OpenCL 2.1 LINUX   Platform Profile                                FULL_PROFILE   Platform Extensions                            cl_khr_icd cl_khr_global_int32_base_atomics cl_khr_global_int32_extended_atomics cl_khr_local_int32_base_atomics cl_khr_local_int32_extended_atomics cl_khr_byte_addressable_store cl_khr_depth_images cl_khr_3d_image_writes cl_intel_exec_by_local_thread cl_khr_spir cl_khr_fp64 cl_khr_image2d_from_buffer cl_intel_vec_len_hint   Platform Host timer resolution                  1ns   Platform Extensions function suffix            INTEL   Platform Name                                  Portable Computing Language Number of devices                                1   Device Name                                    pthread-Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz   Device Vendor                                  GenuineIntel   Device Vendor ID                                0x6c636f70   Device Version                                  OpenCL 1.2 pocl HSTR: pthread-x86_64-pc-linux-gnu-skylake   Driver Version                                  1.4   Device OpenCL C Version                        OpenCL C 1.2 pocl   Device Type                                    CPU   Device Profile                                  FULL_PROFILE   Device Available                                Yes   Compiler Available                              Yes   Linker Available                                Yes   Max compute units                              8   Max clock frequency                            4500MHz   Device Partition                                (core)     Max number of sub-devices                    8     Supported partition types                    equally, by counts     Supported affinity domains                    (n/a)   Max work item dimensions                        3   Max work item sizes                            4096x4096x4096   Max work group size                            4096   Preferred work group size multiple              8   Preferred / native vector sizes                  char                                                16 / 16        short                                              16 / 16        int                                                  8 / 8        long                                                4 / 4        half                                                0 / 0        (n/a)     float                                                8 / 8        double                                              4 / 4        (cl_khr_fp64)   Half-precision Floating-point support          (n/a)   Single-precision Floating-point support        (core)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                Yes     Round to infinity                            Yes     IEEE754-2008 fused multiply-add              Yes     Support is emulated in software              No     Correctly-rounded divide and sqrt operations  Yes   Double-precision Floating-point support        (cl_khr_fp64)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                Yes     Round to infinity                            Yes     IEEE754-2008 fused multiply-add              Yes     Support is emulated in software              No   Address bits                                    64, Little-Endian   Global memory size                              31344009216 (29.19GiB)   Error Correction support                        No   Max memory allocation                          8589934592 (8GiB)   Unified memory for Host and Device              Yes   Minimum alignment for any data type            128 bytes   Alignment of base address                      1024 bits (128 bytes)   Global Memory cache type                        Read/Write   Global Memory cache size                        8388608 (8MiB)   Global Memory cache line size                  64 bytes   Image support                                  Yes     Max number of samplers per kernel            16     Max size for 1D images from buffer            536870912 pixels     Max 1D or 2D image array size                2048 images     Max 2D image size                            16384x16384 pixels     Max 3D image size                            2048x2048x2048 pixels     Max number of read image args                128     Max number of write image args                128   Local memory type                              Global   Local memory size                              4194304 (4MiB)   Max number of constant args                    8   Max constant buffer size                        4194304 (4MiB)   Max size of kernel argument                    1024   Queue properties                                  Out-of-order execution                        Yes     Profiling                                    Yes   Prefer user sync for interop                    Yes   Profiling timer resolution                      1ns   Execution capabilities                            Run OpenCL kernels                            Yes     Run native kernels                            Yes   printf() buffer size                            16777216 (16MiB)   Built-in kernels                                (n/a)   Device Extensions                              cl_khr_byte_addressable_store cl_khr_global_int32_base_atomics cl_khr_global_int32_extended_atomics cl_khr_local_int32_base_atomics cl_khr_local_int32_extended_atomics cl_khr_3d_image_writes cl_khr_fp64 cl_khr_int64_base_atomics cl_khr_int64_extended_atomics cl_khr_fp64   Platform Name                                  NVIDIA CUDA Number of devices                                1   Device Name                                    NVIDIA GeForce GTX 1650   Device Vendor                                  NVIDIA Corporation   Device Vendor ID                                0x10de   Device Version                                  OpenCL 3.0 CUDA   Driver Version                                  495.44   Device OpenCL C Version                        OpenCL C 1.2   Device Type                                    GPU   Device Topology (NV)                            PCI-E, 01:00.0   Device Profile                                  FULL_PROFILE   Device Available                                Yes   Compiler Available                              Yes   Linker Available                                Yes   Max compute units                              14   Max clock frequency                            1515MHz   Compute Capability (NV)                        7.5   Device Partition                                (core)     Max number of sub-devices                    1     Supported partition types                    None     Supported affinity domains                    (n/a)   Max work item dimensions                        3   Max work item sizes                            1024x1024x64   Max work group size                            1024   Preferred work group size multiple              32   Warp size (NV)                                  32   Max sub-groups per work group                  0   Preferred / native vector sizes                  char                                                1 / 1        short                                                1 / 1        int                                                  1 / 1        long                                                1 / 1        half                                                0 / 0        (n/a)     float                                                1 / 1        double                                              1 / 1        (cl_khr_fp64)   Half-precision Floating-point support          (n/a)   Single-precision Floating-point support        (core)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                Yes     Round to infinity                            Yes     IEEE754-2008 fused multiply-add              Yes     Support is emulated in software              No     Correctly-rounded divide and sqrt operations  Yes   Double-precision Floating-point support        (cl_khr_fp64)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                Yes     Round to infinity                            Yes     IEEE754-2008 fused multiply-add              Yes     Support is emulated in software              No   Address bits                                    64, Little-Endian   Global memory size                              4101898240 (3.82GiB)   Error Correction support                        No   Max memory allocation                          1025474560 (978MiB)   Unified memory for Host and Device              No   Integrated memory (NV)                          No   Shared Virtual Memory (SVM) capabilities        (core)     Coarse-grained buffer sharing                Yes     Fine-grained buffer sharing                  No     Fine-grained system sharing                  No     Atomics                                      No   Minimum alignment for any data type            128 bytes   Alignment of base address                      4096 bits (512 bytes)   Preferred alignment for atomics                  SVM                                          0 bytes     Global                                        0 bytes     Local                                        0 bytes   Max size for global variable                    0   Preferred total size of global vars            0   Global Memory cache type                        Read/Write   Global Memory cache size                        458752 (448KiB)   Global Memory cache line size                  128 bytes   Image support                                  Yes     Max number of samplers per kernel            32     Max size for 1D images from buffer            268435456 pixels     Max 1D or 2D image array size                2048 images     Max 2D image size                            32768x32768 pixels     Max 3D image size                            16384x16384x16384 pixels     Max number of read image args                256     Max number of write image args                32     Max number of read/write image args          0   Max number of pipe args                        0   Max active pipe reservations                    0   Max pipe packet size                            0   Local memory type                              Local   Local memory size                              49152 (48KiB)   Registers per block (NV)                        65536   Max number of constant args                    9   Max constant buffer size                        65536 (64KiB)   Max size of kernel argument                    4352 (4.25KiB)   Queue properties (on host)                        Out-of-order execution                        Yes     Profiling                                    Yes   Queue properties (on device)                      Out-of-order execution                        No     Profiling                                    No     Preferred size                                0     Max size                                      0   Max queues on device                            0   Max events on device                            0   Prefer user sync for interop                    No   Profiling timer resolution                      1000ns   Execution capabilities                            Run OpenCL kernels                            Yes     Run native kernels                            No     Sub-group independent forward progress        No     Kernel execution timeout (NV)                Yes   Concurrent copy and kernel execution (NV)      Yes     Number of async copy engines                  3     IL version                                    (n/a)   printf() buffer size                            1048576 (1024KiB)   Built-in kernels                                (n/a)   Device Extensions                              cl_khr_global_int32_base_atomics cl_khr_global_int32_extended_atomics cl_khr_local_int32_base_atomics cl_khr_local_int32_extended_atomics cl_khr_fp64 cl_khr_3d_image_writes cl_khr_byte_addressable_store cl_khr_icd cl_khr_gl_sharing cl_nv_compiler_options cl_nv_device_attribute_query cl_nv_pragma_unroll cl_nv_copy_opts cl_nv_create_buffer cl_khr_int64_base_atomics cl_khr_int64_extended_atomics cl_khr_device_uuid cl_khr_pci_bus_info   Platform Name                                  Intel(R) CPU Runtime for OpenCL(TM) Applications Number of devices                                1   Device Name                                    Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz   Device Vendor                                  Intel(R) Corporation   Device Vendor ID                                0x8086   Device Version                                  OpenCL 2.1 (Build 0)   Driver Version                                  18.1.0.0920   Device OpenCL C Version                        OpenCL C 2.0   Device Type                                    CPU   Device Profile                                  FULL_PROFILE   Device Available                                Yes   Compiler Available                              Yes   Linker Available                                Yes   Max compute units                              8   Max clock frequency                            2500MHz   Device Partition                                (core)     Max number of sub-devices                    8     Supported partition types                    by counts, equally, by names (Intel)     Supported affinity domains                    (n/a)   Max work item dimensions                        3   Max work item sizes                            8192x8192x8192   Max work group size                            8192   Preferred work group size multiple              128   Max sub-groups per work group                  1   Preferred / native vector sizes                  char                                                1 / 32        short                                                1 / 16        int                                                  1 / 8        long                                                1 / 4        half                                                0 / 0        (n/a)     float                                                1 / 8        double                                              1 / 4        (cl_khr_fp64)   Half-precision Floating-point support          (n/a)   Single-precision Floating-point support        (core)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                No     Round to infinity                            No     IEEE754-2008 fused multiply-add              No     Support is emulated in software              No     Correctly-rounded divide and sqrt operations  No   Double-precision Floating-point support        (cl_khr_fp64)     Denormals                                    Yes     Infinity and NANs                            Yes     Round to nearest                              Yes     Round to zero                                Yes     Round to infinity                            Yes     IEEE754-2008 fused multiply-add              Yes     Support is emulated in software              No   Address bits                                    64, Little-Endian   Global memory size                              33491492864 (31.19GiB)   Error Correction support                        No   Max memory allocation                          8372873216 (7.798GiB)   Unified memory for Host and Device              Yes   Shared Virtual Memory (SVM) capabilities        (core)     Coarse-grained buffer sharing                Yes     Fine-grained buffer sharing                  Yes     Fine-grained system sharing                  Yes     Atomics                                      Yes   Minimum alignment for any data type            128 bytes   Alignment of base address                      1024 bits (128 bytes)   Preferred alignment for atomics                  SVM                                          64 bytes     Global                                        64 bytes     Local                                        0 bytes   Max size for global variable                    65536 (64KiB)   Preferred total size of global vars            65536 (64KiB)   Global Memory cache type                        Read/Write   Global Memory cache size                        262144 (256KiB)   Global Memory cache line size                  64 bytes   Image support                                  Yes     Max number of samplers per kernel            480     Max size for 1D images from buffer            523304576 pixels     Max 1D or 2D image array size                2048 images     Base address alignment for 2D image buffers  64 bytes     Pitch alignment for 2D image buffers          64 pixels     Max 2D image size                            16384x16384 pixels     Max 3D image size                            2048x2048x2048 pixels     Max number of read image args                480     Max number of write image args                480     Max number of read/write image args          480   Max number of pipe args                        16   Max active pipe reservations                    32767   Max pipe packet size                            1024   Local memory type                              Global   Local memory size                              32768 (32KiB)   Max number of constant args                    480   Max constant buffer size                        131072 (128KiB)   Max size of kernel argument                    3840 (3.75KiB)   Queue properties (on host)                        Out-of-order execution                        Yes     Profiling                                    Yes     Local thread execution (Intel)                Yes   Queue properties (on device)                      Out-of-order execution                        Yes     Profiling                                    Yes     Preferred size                                4294967295 (4GiB)     Max size                                      4294967295 (4GiB)   Max queues on device                            4294967295   Max events on device                            4294967295   Prefer user sync for interop                    No   Profiling timer resolution                      1ns   Execution capabilities                            Run OpenCL kernels                            Yes     Run native kernels                            Yes     Sub-group independent forward progress        No     IL version                                    SPIR-V_1.0     SPIR versions                                1.2   printf() buffer size                            1048576 (1024KiB)   Built-in kernels                                (n/a)   Device Extensions                              cl_khr_icd cl_khr_global_int32_base_atomics cl_khr_global_int32_extended_atomics cl_khr_local_int32_base_atomics cl_khr_local_int32_extended_atomics cl_khr_byte_addressable_store cl_khr_depth_images cl_khr_3d_image_writes cl_intel_exec_by_local_thread cl_khr_spir cl_khr_fp64 cl_khr_image2d_from_buffer cl_intel_vec_len_hint NULL platform behavior   clGetPlatformInfo(NULL, CL_PLATFORM_NAME, ...)  No platform   clGetDeviceIDs(NULL, CL_DEVICE_TYPE_ALL, ...)  No platform   clCreateContext(NULL, ...) [default]            No platform   clCreateContext(NULL, ...) [other]              Success [POCL]   clCreateContextFromType(NULL, CL_DEVICE_TYPE_DEFAULT)  No platform   clCreateContextFromType(NULL, CL_DEVICE_TYPE_CPU)  No platform   clCreateContextFromType(NULL, CL_DEVICE_TYPE_GPU)  No platform   clCreateContextFromType(NULL, CL_DEVICE_TYPE_ACCELERATOR)  No platform   clCreateContextFromType(NULL, CL_DEVICE_TYPE_CUSTOM)  No platform   clCreateContextFromType(NULL, CL_DEVICE_TYPE_ALL)  No platform NOTE: your OpenCL library only supports OpenCL 2.1, but some installed platforms support OpenCL 3.0. Programs using 3.0 features may crash or behave unexpectedly


vga adapters
Code:
sudo lspci -v | grep -i vga 00:02.0 VGA compatible controller: Intel Corporation UHD Graphics (rev 05) (prog-if 00 [VGA controller]) 01:00.0 VGA compatible controller: NVIDIA Corporation Device 1f99 (rev a1) (prog-if 00 [VGA controller])


test(intel(skipped) +nvidia)
Code:
root@pc:/home/tester# hashcat -b -m 0 -D 3,2 hashcat (v5.1.0) starting in benchmark mode... Benchmarking uses hand-optimized kernel code by default. You can use it in your cracking session by setting the -O option. Note: Using optimized kernel code limits the maximum supported password length. To disable the optimized kernel code in benchmark mode, use the -w option. * Device #1: Not a native Intel OpenCL runtime. Expect massive speed loss.             You can use --force to override, but do not report related errors. * Device #2: WARNING! Kernel exec timeout is not disabled.             This may cause "CL_OUT_OF_RESOURCES" or related errors.             To disable the timeout, see: https://hashcat.net/q/timeoutpatch nvmlDeviceGetFanSpeed(): Not Supported OpenCL Platform #1: The pocl project ==================================== * Device #1: pthread-Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz, skipped. OpenCL Platform #2: NVIDIA Corporation ====================================== * Device #2: NVIDIA GeForce GTX 1650, 977/3911 MB allocatable, 14MCU OpenCL Platform #3: Intel(R) Corporation ======================================== * Device #3: Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz, skipped. Benchmark relevant options: =========================== * --opencl-device-types=3,2 * --optimized-kernel-enable Hashmode: 0 - MD5 Speed.#2.........:  8995.5 MH/s (51.83ms) @ Accel:512 Loops:256 Thr:256 Vec:1 Started: Sun Dec  5 11:36:44 2021 Stopped: Sun Dec  5 11:36:51 2021


Why it skipped?
Why does hashcat not want to use the third device?



test(pocl + nvidia)
Code:
root@pc:/home/tester# hashcat -b -m 0 -D 1,2 hashcat (v5.1.0) starting in benchmark mode... Benchmarking uses hand-optimized kernel code by default. You can use it in your cracking session by setting the -O option. Note: Using optimized kernel code limits the maximum supported password length. To disable the optimized kernel code in benchmark mode, use the -w option. * Device #1: Not a native Intel OpenCL runtime. Expect massive speed loss.             You can use --force to override, but do not report related errors. * Device #2: WARNING! Kernel exec timeout is not disabled.             This may cause "CL_OUT_OF_RESOURCES" or related errors.             To disable the timeout, see: https://hashcat.net/q/timeoutpatch nvmlDeviceGetFanSpeed(): Not Supported OpenCL Platform #1: The pocl project ==================================== * Device #1: pthread-Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz, skipped. OpenCL Platform #2: NVIDIA Corporation ====================================== * Device #2: NVIDIA GeForce GTX 1650, 977/3911 MB allocatable, 14MCU OpenCL Platform #3: Intel(R) Corporation ======================================== * Device #3: Intel(R) Core(TM) i5-10300H CPU @ 2.50GHz, 7984/31939 MB allocatable, 8MCU Benchmark relevant options: =========================== * --opencl-device-types=1,2 * --optimized-kernel-enable Hashmode: 0 - MD5 Speed.#2.........:  8522.4 MH/s (51.04ms) @ Accel:512 Loops:256 Thr:256 Vec:1 Speed.#3.........:  323.6 MH/s (27.29ms) @ Accel:1024 Loops:1024 Thr:1 Vec:8 Speed.#*.........:  8846.1 MH/s Started: Sun Dec  5 11:37:35 2021 Stopped: Sun Dec  5 11:37:43 2021


PS:
Сonfig xorg.conf was not found in my system:
Code:
root@pc:/home/tester# find / -iname *xorg.conf* find: ‘/run/user/1000/gvfs’: Permission denied find: ‘/run/user/125/gvfs’: Permission denied /usr/share/man/man5/xorg.conf.d.5.gz /usr/share/man/man5/xorg.conf.5.gz /usr/share/X11/xorg.conf.d /usr/share/doc/xserver-xorg-video-intel/xorg.conf /etc/X11/xorg.conf.nvidia-xconfig-original # this is the correct config ??? cat /usr/share/doc/xserver-xorg-video-intel/xorg.conf Section "Device" Identifier "Intel" Driver "intel" # Option "AccelMethod" "uxa" EndSection



Code:
Newer versions of Ubuntu (and maybe other distributions as well) do not use xorg.conf any longer. Alternatively they use the folder /usr/share/x11/xorg.conf.d/ where you can put in snippets of the config. Create a new file, call it “20-nvidia.conf” and put in the following content: Section "Device" Identifier "MyGPU" Driver "nvidia" Option "Interactive" "0" EndSection
if I create this file, my system does not boot at all, it freezes on the logo.
Reply
#2
Firstly,  using case sensitive parameters is your issue. There is a difference between -d and -D. -D will define to use different OpenCL device types (CPU or GPU) and -d will determine to use device #'s. So saying -d 2,3 is what you were intended to use rather than -D is my guess.

Second, using old versions of hashcat could also cause problems. Just download the latest version, no reason to have such an old version. I'm guessing you're reading some aged tutorial.

Third, POCL project doesn't co-operate very well with hashcat from the several 100's of post I have read about having issues. Just download the proper drivers for your hardware and operating system. 

The homepage gives you a great idea of the necessary components for proper operation of hashcat. 
https://hashcat.net/hashcat/
Reply
#3
(12-06-2021, 05:09 AM)slyexe Wrote: Firstly,  using case sensitive parameters is your issue. There is a difference between -d and -D. -D will define to use different OpenCL device types (CPU or GPU) and -d will determine to use device #'s. So saying -d 2,3 is what you were intended to use rather than -D is my guess.

Second, using old versions of hashcat could also cause problems. Just download the latest version, no reason to have such an old version. I'm guessing you're reading some aged tutorial.

Third, POCL project doesn't co-operate very well with hashcat from the several 100's of post I have read about having issues. Just download the proper drivers for your hardware and operating system. 

The homepage gives you a great idea of the necessary components for proper operation of hashcat. 
https://hashcat.net/hashcat/

1. In the latest ubuntu from the repository, only this version is available.
2. When i download latest hashcat(binary) and placing it in "/ bin", my video card disappears from my devices.

Thanks for the answer.
Reply