Support and discussions for Molcas and OpenMolcas users and developers
You are not logged in.
Hello there!
I'm using OpenMolcas v. 26.02 and I'm now trying to optimize conical intersection geometries for a 24 atoms system with SA-CASSCF method, but a sudden stop in module ALASKA occurred.
Input orbitals are from a previously converged SA-CASSCF(12,11) single-point for the starting geometry which is a candidate for S1/S0 crossing geometry.
Here's the input
&GATEWAY
coord=v_oriri.xyz
basis=def2-tzvp
group=nosym
Constraints
a=Ediff 1 2
Value
a=0.000
End of Constraints
>> DO WHILE <<
&SEWARD
DoAnalytical
CHOInput
1-CEnter
EndCHOInput
&RASSCF
FILEOrb=cas_oriri.RasOrb
spin=1
nactel=12
inactive=54
ras2=11
CIROOT=2 2 1
RLXRoot=1
&SLAPAF
ITERATIONS=300
>> ENDDO <<and I got this non-zero code exit at the output after convergence was signaled after 18 cycles
--- Stop Module: mclr at Fri Apr 10 17:42:50 2026 /rc=_RC_ALL_IS_WELL_ ---
*** files: xmldump
saved to directory $SCRIPT_SUBMIT_DIR
--- Module mclr spent 23 hours 1 minute 41 seconds ---
*** symbolic link created: INPORB -> opt_oriri.RasOrb
--- Start Module: alaska at Fri Apr 10 17:43:26 2026 ---
()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()
&ALASKA
launched 16 MPI processes, running in PARALLEL mode (work-sharing enabled)
available to each process: 10 GB of memory, 1 thread
master pid: 1026832
()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()()
Threshold for contributions to the gradient: 1.000E-07
********************************************
* Symmetry Adapted Cartesian Displacements *
********************************************
Irreducible representation : a
Basis function(s) of irrep: x, y, xy, Rz, z, xz, Ry, yz, Rx, I
Basis Label Type Center Phase
1 N1 x 1 1
2 N1 y 1 1
3 N1 z 1 1
4 C2 x 2 1
5 C2 y 2 1
6 C2 z 2 1
7 C3 x 3 1
8 C3 y 3 1
9 C3 z 3 1
10 N4 x 4 1
11 N4 y 4 1
12 N4 z 4 1
13 C5 x 5 1
14 C5 y 5 1
15 C5 z 5 1
16 C6 x 6 1
17 C6 y 6 1
18 C6 z 6 1
19 N7 x 7 1
20 N7 y 7 1
21 N7 z 7 1
22 C8 x 8 1
23 C8 y 8 1
24 C8 z 8 1
25 N9 x 9 1
26 N9 y 9 1
27 N9 z 9 1
28 C10 x 10 1
29 C10 y 10 1
30 C10 z 10 1
31 N11 x 11 1
32 N11 y 11 1
33 N11 z 11 1
34 C12 x 12 1
35 C12 y 12 1
36 C12 z 12 1
37 C13 x 13 1
38 C13 y 13 1
39 C13 z 13 1
40 S14 x 14 1
41 S14 y 14 1
42 S14 z 14 1
43 N15 x 15 1
44 N15 y 15 1
45 N15 z 15 1
46 C16 x 16 1
47 C16 y 16 1
48 C16 z 16 1
49 H17 x 17 1
50 H17 y 17 1
51 H17 z 17 1
52 H18 x 18 1
53 H18 y 18 1
54 H18 z 18 1
55 H19 x 19 1
56 H19 y 19 1
57 H19 z 19 1
58 H20 x 20 1
59 H20 y 20 1
60 H20 z 20 1
61 H21 x 21 1
62 H21 y 21 1
63 H21 z 21 1
64 H22 x 22 1
65 H22 y 22 1
66 H22 z 22 1
67 H23 x 23 1
68 H23 y 23 1
69 H23 z 23 1
70 H24 x 24 1
71 H24 y 24 1
72 H24 z 24 1
No automatic utilization of translational and rotational invariance of the energy is employed.
Cholesky-ERI gradients!
--- Stop Module: alaska at Fri Apr 10 17:45:24 2026 /rc=_RC_EXTERNAL_TERM_ ---
*** files: xmldump
saved to directory $SCRIPT_SUBMIT_DIR
--- Module alaska spent 1 minute 58 seconds ---
.########################.
.# Non-zero return code #.
.########################.
Timing: Wall=84046.14 User=37008.39 System=6989.19I looked for similar problems in forum and it didn't helped me to figure out what is wrong. Thank you.
Last edited by tuancpacheco (2026-04-13 11:17:46)
Offline
I thought it could be some bug or memory issue from the cluster because of the return code _EXTERNAL_TERM_, then I launched 32 processors instead of 16 for the same single job with 10000 MB for each and the same crash occurred. What else could it be?
The SLURM output ends up with
--------------------------------------------------------------------------
Primary job terminated normally, but 1 process returned
a non-zero exit code. Per user-direction, the job has been aborted.
--------------------------------------------------------------------------
--------------------------------------------------------------------------
mpiexec noticed that process rank 25 with PID 1066623 on node n02 exited on signal 9 (Killed).
--------------------------------------------------------------------------
slurmstepd-n02: error: Detected 8 oom_kill events in StepId=550.batch. Some of the step tasks have been OOM Killed.Last edited by tuancpacheco (2026-04-13 11:25:41)
Offline
oom_kill means "out of memory" (I believe). Make sure *each* of the processes has access to the requested 10GB (non-overlapping). More processes will only make it more likely to crash.
Offline
Thanks Ignacio. I launched 2 nodes (16 processors each) and reduced mem to 8000 MB. Now the return code is not external nor oom_kill anymore. It crashes at the same step but shows
No automatic utilization of translational and rotational invariance of the energy is employed.
Cholesky-ERI gradients!
Not Implemented for Cholesky yet!
--- Stop Module: alaska at Tue Apr 14 11:37:19 2026 /rc=41 ---
*** files: xmldump
saved to directory $SCRIPT_SUBMIT_DIR
--- Module alaska spent 10 minutes 31 seconds ---
.########################.
.# Non-zero return code #.
.########################.
Timing: Wall=4391.09 User=24712.47 System=15124.19I thus launched with no Cholesky keyword and it is too slow, apparently spending more than 13h in first Seward module 
Last edited by tuancpacheco (2026-04-16 02:31:57)
Offline
Remove ChoInput and add RICD in GATEWAY.
Offline
It worked!!! Crystal Clear convergence!
Thanks a lot, Ignacio! My investigation can proceed hereafter.
Offline