联系:手机/微信(+86 17813235971) QQ(107644445)
标题:公有云安装19c rac遇到问题—169网段udp异常
作者:惜分飞©版权所有[未经本人同意,不得以任何形式转载,否则有进一步追究法律责任的权利.]
应客户要求在xx公有云上面安装19c rac,通过各方的努力,最后安装情况如下
1. 两个节点root.sh执行成功,crs启动正常,asm磁盘组访问正常,但是有一个节点asm实例无法启动,一个节点的db实例无法启动
---节点1 [root@dzbl1 ~]# su - grid Last login: Thu May 20 12:32:55 CST 2021 [grid@dzbl1 ~]$ ps -ef|grep ASM grid 477 1 0 May19 ? 00:00:24 /u01/app/19c/grid/bin/tnslsnr ASMNET1LSNR_ASM -no_crs_notify -inherit grid 22075 22039 0 12:42 pts/1 00:00:00 grep --color=auto ASM [grid@dzbl1 ~]$ asmcmd ASMCMD> lsdg State Type Rebal Sector Logical_Sector Block AU Total_MB Free_MB Req_mir_free_MB Usable_file_MB Offline_disks Voting_files Name MOUNTED EXTERN N 512 512 4096 4194304 1907344 1904420 0 1904420 0 N DATA/ MOUNTED EXTERN N 512 512 4096 4194304 1150344 1149032 0 1149032 0 N FRA/ MOUNTED EXTERN N 512 512 4096 4194304 14304 13988 0 13988 0 Y SYSTEMDG/ ASMCMD> exit [grid@dzbl1 ~]$ crsctl status res -t -------------------------------------------------------------------------------- Name Target State Server State details -------------------------------------------------------------------------------- Local Resources -------------------------------------------------------------------------------- ora.LISTENER.lsnr ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.chad ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.net1.network ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.ons ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.proxy_advm OFFLINE OFFLINE dzbl1 STABLE OFFLINE OFFLINE dzbl2 STABLE -------------------------------------------------------------------------------- Cluster Resources -------------------------------------------------------------------------------- ora.ASMNET1LSNR_ASM.lsnr(ora.asmgroup) 1 ONLINE ONLINE dzbl1 STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 ONLINE OFFLINE STABLE ora.DATA.dg(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.FRA.dg(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.LISTENER_SCAN1.lsnr 1 ONLINE ONLINE dzbl2 STABLE ora.SYSTEMDG.dg(ora.asmgroup) 1 OFFLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.asm(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 Started,STABLE 3 OFFLINE OFFLINE STABLE ora.asmnet1.asmnetwork(ora.asmgroup) 1 ONLINE ONLINE dzbl1 STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.cvu 1 ONLINE ONLINE dzbl2 STABLE ora.dzbl1.vip 1 ONLINE ONLINE dzbl1 STABLE ora.dzbl2.vip 1 ONLINE ONLINE dzbl2 STABLE ora.dzbldb.db 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 Open,HOME=/u01/app/o racle/product/19c/db _1,STABLE ora.qosmserver 1 ONLINE ONLINE dzbl2 STABLE ora.scan1.vip 1 ONLINE ONLINE dzbl2 STABLE -------------------------------------------------------------------------------- [grid@dzbl1 ~]$ ---节点2 [grid@dzbl2 ~]$ ps -ef|grep ASM grid 2464 1 0 May18 ? 00:00:29 /u01/app/19c/grid/bin/tnslsnr ASMNET1LSNR_ASM -no_crs_notify -inherit grid 6826 1 0 May19 ? 00:00:09 oracle+ASM2_asmb_dzbldb2 (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) grid 14089 1 0 12:38 ? 00:00:00 asm_m000_+ASM2 grid 15670 1 0 12:40 ? 00:00:00 oracle+ASM2_crf (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) grid 16503 1 0 May18 ? 00:00:05 asm_pmon_+ASM2 grid 16505 1 0 May18 ? 00:00:04 asm_clmn_+ASM2 grid 16507 1 0 May18 ? 00:00:11 asm_psp0_+ASM2 grid 16518 1 0 12:42 ? 00:00:00 oracle+ASM2 (LOCAL=NO) grid 16562 1 0 May18 ? 00:18:22 asm_vktm_+ASM2 grid 16567 1 0 May18 ? 00:00:08 asm_gen0_+ASM2 grid 16569 1 0 May18 ? 00:00:02 asm_mman_+ASM2 grid 16573 1 0 May18 ? 00:00:06 asm_gen1_+ASM2 grid 16577 1 0 May18 ? 00:01:13 asm_diag_+ASM2 grid 16579 1 0 May18 ? 00:00:04 asm_ping_+ASM2 grid 16581 1 0 May18 ? 00:00:09 asm_pman_+ASM2 grid 16583 1 0 May18 ? 00:03:08 asm_dia0_+ASM2 grid 16585 1 0 May18 ? 00:01:41 asm_lmon_+ASM2 grid 16587 1 0 May18 ? 00:01:55 asm_lmd0_+ASM2 grid 16589 1 0 May18 ? 00:04:26 asm_lms0_+ASM2 grid 16591 1 0 May18 ? 00:02:13 asm_lmhb_+ASM2 grid 16596 1 0 May18 ? 00:00:02 asm_lck1_+ASM2 grid 16598 1 0 May18 ? 00:00:02 asm_dbw0_+ASM2 grid 16600 1 0 May18 ? 00:00:02 asm_lgwr_+ASM2 grid 16602 1 0 May18 ? 00:00:05 asm_ckpt_+ASM2 grid 16604 1 0 May18 ? 00:00:01 asm_smon_+ASM2 grid 16606 1 0 May18 ? 00:00:02 asm_lreg_+ASM2 grid 16608 1 0 May18 ? 00:00:01 asm_pxmn_+ASM2 grid 16610 1 0 May18 ? 00:00:11 asm_rbal_+ASM2 grid 16612 1 0 May18 ? 00:00:24 asm_gmon_+ASM2 grid 16614 1 0 May18 ? 00:00:06 asm_mmon_+ASM2 grid 16616 1 0 May18 ? 00:00:47 asm_mmnl_+ASM2 grid 16618 1 0 May18 ? 00:02:52 asm_imr0_+ASM2 grid 16627 1 0 May18 ? 00:00:30 asm_scm0_+ASM2 grid 16633 1 0 May18 ? 00:00:11 asm_lck0_+ASM2 grid 16662 1 0 May18 ? 00:07:10 asm_gcr0_+ASM2 grid 16699 1 0 May19 ? 00:00:00 oracle+ASM2 (LOCAL=NO) grid 16746 1 0 May18 ? 00:00:06 asm_asmb_+ASM2 grid 16748 1 0 May18 ? 00:00:13 oracle+ASM2_asmb_+asm2 (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) grid 16756 1 0 May18 ? 00:00:00 oracle+ASM2_ocr (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) grid 17567 1 0 May18 ? 00:00:00 oracle+ASM2 (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) grid 17622 17536 0 12:43 pts/1 00:00:00 grep --color=auto ASM grid 27829 1 0 May18 ? 00:00:00 oracle+ASM2 (DESCRIPTION=(LOCAL=YES)(ADDRESS=(PROTOCOL=beq))) [grid@dzbl2 ~]$ asmcmd ASMCMD> lsdg State Type Rebal Sector Logical_Sector Block AU Total_MB Free_MB Req_mir_free_MB Usable_file_MB Offline_disks Voting_files Name MOUNTED EXTERN N 512 512 4096 4194304 1907344 1904420 0 1904420 0 N DATA/ MOUNTED EXTERN N 512 512 4096 4194304 1150344 1149032 0 1149032 0 N FRA/ MOUNTED EXTERN N 512 512 4096 4194304 14304 13988 0 13988 0 Y SYSTEMDG/ ASMCMD> exit [grid@dzbl2 ~]$ crsctl stat res -t -------------------------------------------------------------------------------- Name Target State Server State details -------------------------------------------------------------------------------- Local Resources -------------------------------------------------------------------------------- ora.LISTENER.lsnr ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.chad ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.net1.network ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.ons ONLINE ONLINE dzbl1 STABLE ONLINE ONLINE dzbl2 STABLE ora.proxy_advm OFFLINE OFFLINE dzbl1 STABLE OFFLINE OFFLINE dzbl2 STABLE -------------------------------------------------------------------------------- Cluster Resources -------------------------------------------------------------------------------- ora.ASMNET1LSNR_ASM.lsnr(ora.asmgroup) 1 ONLINE ONLINE dzbl1 STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 ONLINE OFFLINE STABLE ora.DATA.dg(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.FRA.dg(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.LISTENER_SCAN1.lsnr 1 ONLINE ONLINE dzbl2 STABLE ora.SYSTEMDG.dg(ora.asmgroup) 1 OFFLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.asm(ora.asmgroup) 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 Started,STABLE 3 OFFLINE OFFLINE STABLE ora.asmnet1.asmnetwork(ora.asmgroup) 1 ONLINE ONLINE dzbl1 STABLE 2 ONLINE ONLINE dzbl2 STABLE 3 OFFLINE OFFLINE STABLE ora.cvu 1 ONLINE ONLINE dzbl2 STABLE ora.dzbl1.vip 1 ONLINE ONLINE dzbl1 STABLE ora.dzbl2.vip 1 ONLINE ONLINE dzbl2 STABLE ora.dzbldb.db 1 ONLINE OFFLINE STABLE 2 ONLINE ONLINE dzbl2 Open,HOME=/u01/app/o racle/product/19c/db _1,STABLE ora.qosmserver 1 ONLINE ONLINE dzbl2 STABLE ora.scan1.vip 1 ONLINE ONLINE dzbl2 STABLE -------------------------------------------------------------------------------- [grid@dzbl2 ~]$
2. 分析db和asm有一个实例无法启动原因分析
--实例启动报错 SQL> startup ORA-03113: end-of-file on communication channel --无法启动节点alert日志 2021-05-19T12:41:32.143124+08:00 NOTE: ASMB (index:0) registering with ASM instance as Flex client 0xffffffffffffffff (reg:2449521867) (startid:1072960888) (new connection) 2021-05-19T12:41:32.349766+08:00 My CSS node number is 1 My CSS hostname is dzbl1 lmon registered with NM - instance number 1 (internal mem no 0) 2021-05-19T12:41:34.054865+08:00 Using default pga_aggregate_limit of 16384 MB 2021-05-19T12:42:16.978085+08:00 No connectivity to other instances in the cluster during startup. Hence, LMON is terminating the instance. Please check the LMON trace file for details. Also, please check the network logs of this instance along with clusterwide network health for problems and then re-start this instance. LMON (ospid: ): terminating the instance due to ORA error Cause - 'Instance is being terminated by LMON' 2021-05-19T12:42:17.115807+08:00 System state dump requested by (instance=1, osid=29660 (LMON)), summary=[abnormal instance termination]. error - 'Instance is terminating. System State dumped to trace file /u01/app/oracle/diag/rdbms/dzbldb/dzbldb1/trace/dzbldb1_diag_29641.trc 2021-05-19T12:42:17.227469+08:00 Dumping diagnostic data in directory=[cdmp_20210519124217], requested by (instance=1, osid=29660 (LMON)), summary=[abnormal instance termination]. 2021-05-19T12:42:18.344481+08:00 Instance terminated by LMON, pid = 29660 --正常节点lmon日志 *** 2021-05-19T12:42:29.348455+08:00 IPCLW:[0.16]{-}[CNCT]:PROTO: [1621399349248289]Warning! ACNH://0x7f3d993a7990/peer=[UNKNWN]&ospid=0&msn=993097808&seq=995707504 (169.254.14.18:32056) has outstanding sends during delete. IPCLW:[0.17]{-}[CNCT]:UTIL: [1621399349248289] ACNH 0x7f3d993a7990 State: 2 SMSN: 993097806 PKT(993097808.995707504) # Pending: 2 IPCLW:[0.18]{-}[CNCT]:UTIL: [1621399349248289] Peer: [UNKNWN].0 AckSeq: 0 IPCLW:[0.19]{-}[CNCT]:UTIL: [1621399349248289] Flags: 0x40000000 IHint: 0x30693d920000001f THint: 0x0 IPCLW:[0.20]{-}[CNCT]:UTIL: [1621399349248289] Local Address: 169.254.17.231:19443 Remote Address: 169.254.14.18:32056 IPCLW:[0.21]{-}[CNCT]:UTIL: [1621399349248289] Remote PID: ver 0 flags 1 trans 2 tos 0 opts 0 xdata3 165f xdata2 70dbd629 IPCLW:[0.22]{-}[CNCT]:UTIL: [1621399349248289] : mmsz 32768 mmr 4096 mms 4096 xdata c2a71bf9 IPCLW:[0.23]{-}[CNCT]:UTIL: [1621399349248289] IVPort: 46944 TVPort: 7161 IMPT: 25433 RMPT: 5727 Pending Sends: Yes Unacked Sends: Yes IPCLW:[0.24]{-}[CNCT]:UTIL: [1621399349248289] Send Engine Queued: No sshdl -1 ssts 0 rtts 0 snderrchk 0 creqcnt 19 credits 0/0 IPCLW:[0.25]{-}[CNCT]:UTIL: [1621399349248289] Unackd Messages 993097806 -> 993097807. SSEQ 995707502 Send Time: INVALID TIME SMSN # Xmits: 0 EMSN INVALID TIME IPCLW:[0.26]{-}[CNCT]:UTIL: [1621399349248289] Pending send queue: IPCLW:[0.27]{-}[CNCT]:UTIL: [1621399349248289] [0] mbuf 0x7f3d99397770 MSN 993097806 Seq 995707502 -> 995707503 # XMits: 0 IPCLW:[0.28]{-}[CNCT]:UTIL: [1621399349248289] [1] mbuf 0x7f3d99397350 MSN 993097807 Seq 995707503 -> 995707504 # XMits: 0 kjxgfipccb: msg 0x7f3d9934a680, mbo 0x7f3d9934a670, type 24, ack 0, ref 0, stat 34 kjxgfipccb: msg 0x7f3d9934a878, mbo 0x7f3d9934a868, type 18, ack 0, ref 0, stat 34
从日志看异常节点的169.254.14.18:32056和169.254.17.231:19443无法使用udp进行通讯,参考:Only One Instance of a RAC Database Can Start at a Time: Second Instance Fails to Start due to “No reconfig messages from other instances” – LMON is terminating the instance (Doc ID 2528588.1),从而使得asm和db实例只能启动一个节点.到目前为止,初步看很可能是公有云的对于169.254网段的某些限制导致.
对于两个节点asm磁盘组mount,crs正常启动.这个是由于使用的是fiex asm技术实现(在asm实例启动正常情况下直接启动本地asm实例,在本地asm实例无法正常启动,通过fiex asm实现磁盘组正常mount)
公有云那边在策略中把私网的udp限制全部放开,asm和db第二个节点也启动正常