Fedora Account System
Red Hat Associate
Red Hat Customer
Description of problem: ====================== If we try to create a block of size greater than the empty space in volume (with the prealloc=full option), it waits and errors out (or times out) with an unrelated message. "gluster-block list <volname>" however does display the new blockname (even though the creation failed) with all its related info, though the meta information when checked in block-meta displays the field ENTRYCREATE as FAIL. Neither the error message on STDERR, nor the the related logs in /var/log/gluster-block/ hint that the block size exceeds the size of the volume. On the contrary, it leaves a stale block entry in the system- which is not functional. The size checking of the volume should be done even before we start creating a block- thereby we can error out with a logical message at the outset itself. Version-Release number of selected component (if applicable): ============================================================ glusterfs-3.8.4-31 and gluster-block-0.2.1-3 How reproducible: ================ Twice Additional info: =================== Different log messages in the files shown below. The block in question is 'testblock3', and volume name is 'ozone'. gluster_blockd.log ------------------ [2017-07-03 11:19:05.342098] INFO: create cli request, volume=ozone blockname=testblock3 mpath=2 blockhosts=10.70.47.121,10.70.47.113 authmode=0 size=53687091200 [at block_svc_routines.c+1722 :<block_create_cli_1_svc>] [2017-07-03 11:49:09.214215] ERROR: failed while creating block file in gluster volume volume: ozone host: 10.70.47.121,10.70.47.113 [at block_svc_routines.c+1787 :<block_create_cli_1_svc>] gluster-block-cli.log --------------------- [2017-07-03 11:49:09.022376] ERROR: glfs_zerofill(62b60926-7edf-4e07-9f82-eb9cd1b50467): on volume ozone for block testblock3 of size 53687091200 failed[Transport endpoint is not connected] [at glfs-operations.c+140 :<glusterBlockCreateEntry>] gluster-block-gfapi.log ------------------------ [2017-07-03 11:24:05.404699] ERROR: block_create_cli_1: RPC: Timed out block testblock3 create on volume ozone with hosts 10.70.47.121,10.70.47.113 failed [at gluster-block.c+110 :<glusterBlockCliRPC_1>] [2017-07-03 11:24:05.406538] ERROR: failed creating block testblock3 on volume ozone with hosts 10.70.47.121,10.70.47.113 [at gluster-block.c+421 :<glusterBlockCreate>] [2017-07-03 11:24:05.406706] ERROR: failed in create [at gluster-block.c+540 :<glusterBlockParseArgs>] [2017-07-03 11:29:58.349508] ERROR: block_create_cli_1: RPC: Timed out block testblock3 create on volume ozone with hosts 10.70.47.121,10.70.47.113 failed [at gluster-block.c+110 :<glusterBlockCliRPC_1>] [2017-07-03 11:29:58.352169] ERROR: failed creating block testblock3 on volume ozone with hosts 10.70.47.121,10.70.47.113 [at gluster-block.c+421 :<glusterBlockCreate>] [2017-07-03 11:29:58.352229] ERROR: failed in create [at gluster-block.c+540 :<glusterBlockParseArgs>] [2017-07-03 11:49:09.016072] E [rpc-clnt.c:200:call_bail] 0-ozone-client-3: bailing out frame type(GlusterFS 3.3) op(ZEROFILL(46)) xid = 0xeb sent = 2017-07-03 11:19:05.504506. timeout = 1800 for 10.70.47.115:49155 [2017-07-03 11:49:09.016323] W [MSGID: 114031] [client-rpc-fops.c:2109:client3_3_zerofill_cbk] 0-ozone-client-3: remote operation failed [Transport endpoint is not connected]
I have been conveyed that this condition will /never/ be hit in CNS environment as the underlying architecture/process will take care of the size-checking-part. Having said that, this bug does hit the usability aspect of gluster-block feature. I am okay for this bug to be deferred. I would like PM's consensus on the same, hence proposing it as a blocker for it to be discussed in the wider group.
This patch will delete the stale metadata along with the backend storage which is partially created may be due to lack of space or some other error https://review.gluster.org/17716
Since the problem described in this bug report should be resolved in a recent advisory, it has been closed with a resolution of ERRATA. For information on the advisory, and where to find the updated files, follow the link below. If the solution does not work for you, open a new bug report. https://access.redhat.com/errata/RHEA-2018:2691